Skip to main content
leaderboard
.cn
Toggle Sidebar
Search models or organizations…
⌘K
Explore
Frontier Models
Historical Trends
Local Models
Compare Models
Ranking Studio
Library
Models
Pricing
Organizations
Benchmarks
Data sources
Settings
Toggle Sidebar
leaderboard
.cn
Loading page content
Feedback
Frontier Models
Studio
Share
A consolidated ranking based on the evaluations and latest reports in our dataset, scored with Bradley-Terry and Elo pairwise methods.
How the ranking works →
Frontier Models
207
Models
8
Ranked domains
412
Benchmarks
78
Sources
Updated 07/22, 03:19
Overall
410 benchmarks
Combines user preference, general knowledge, and instruction-following ability.
Studio
Model
Score
1
Claude Fable 5
100.0
2
GPT-5.6 Sol
96.8
3
Kimi K3
94.7
4
GPT-5.5
92.6
5
Claude Opus 4.8
92.3
6
GPT-5.6 Terra
90.4
7
Grok 4.5
88.0
8
GLM-5.2
88.0
9
Muse Spark 1.1
87.4
10
Claude Sonnet 5
87.4
Open-weight
248 benchmarks
Includes only open-weight models, combining reasoning, coding, agent, multimodal, and long-context performance.
Studio
Model
Score
1
GLM-5.2
100.0
2
Hunyuan 3
95.3
3
Kimi-K2.7-Code
95.1
Local
107 benchmarks
Frontier models that can run locally and have fewer than 40 billion parameters.
Studio
Model
Score
1
Qwen3.6-27B
100.0
2
Qwen3.6-35B-A3B
93.2
3
Qwen3-30B-A3B-Thinking-2507
89.7
Coding
93 benchmarks
Covers software engineering, competitive programming, and terminal tasks.
Studio
Model
Score
1
Claude Fable 5
100.0
2
GPT-5.6 Sol
97.5
3
Kimi K3
96.8
Research
118 benchmarks
Focuses on scientific knowledge, mathematical reasoning, and research-oriented problem solving.
Studio
Model
Score
1
GPT-5.5
100.0
2
Claude Fable 5
99.5
3
Claude Opus 4.8
99.2
Agents
107 benchmarks
Covers tool use, web browsing, and long-horizon task execution.
Studio
Model
Score
1
Claude Fable 5
100.0
2
GPT-5.6 Sol
98.8
3
Kimi K3
96.0
Vision
81 benchmarks
Covers multimodal understanding, visual mathematics, screen, and video tasks.
Studio
Model
Score
1
Claude Fable 5
100.0
2
Seed2.1 Pro
99.6
3
Kimi K3
98.4
Long context
38 benchmarks
Covers retrieval, synthesis, and reasoning over million-token contexts.
Studio
Model
Score
1
GPT-5.6 Sol
100.0
2
Claude Fable 5
95.8
3
GPT-5.6 Terra
92.6
More domains
More data needed
Domains with limited samples do not show model rankings yet.
View benchmarks
Pending domains
Minimum 3
Writing
2 / 3
A domain joins the homepage ranking automatically after reaching the minimum coverage.