Skip to main content
leaderboard
.cn
Toggle Sidebar
Search models or organizations…
⌘K
Explore
Frontier Models
Historical Trends
Local Models
Compare Models
Ranking Studio
Library
Models
Pricing
Organizations
Benchmarks
Data sources
Settings
Toggle Sidebar
leaderboard
.cn
Loading page content
Feedback
Frontier Models
Studio
Share
A consolidated ranking based on the evaluations and latest reports in our dataset, scored with Bradley-Terry and Elo pairwise methods.
How the ranking works →
Frontier Models
215
Models
8
Ranked domains
470
Benchmarks
85
Sources
Updated 08/12, 01:31
Overall
468 benchmarks
Combines user preference, general knowledge, and instruction-following ability.
Studio
Model
Score
1
Claude Opus 5
100.0
2
Qwen3.8-Max
100.0
3
Claude Fable 5
99.5
4
Kimi K3
98.1
5
GPT-5.6 Sol
97.8
6
GPT-5.5
93.8
7
GPT-5.6 Terra
92.6
8
GPT-5.5 Pro
92.3
9
Claude Opus 4.8
91.4
10
DeepSeek-V4-Flash-0731
90.5
Open-weight
273 benchmarks
Includes only open-weight models, combining reasoning, coding, agent, multimodal, and long-context performance.
Studio
Model
Score
1
Kimi K3
100.0
2
DeepSeek-V4-Flash-0731
99.0
3
GLM-5.2
96.0
Local
122 benchmarks
Frontier models that can run locally and have fewer than 40 billion parameters.
Studio
Model
Score
1
Qwen3.6-27B
100.0
2
Muse Glimmer 30B
95.1
3
Qwen3.5-27B
94.7
Coding
104 benchmarks
Covers software engineering, competitive programming, and terminal tasks.
Studio
Model
Score
1
Claude Fable 5
100.0
2
Claude Opus 5
96.0
3
GPT-5.6 Sol
95.9
Research
131 benchmarks
Focuses on scientific knowledge, mathematical reasoning, and research-oriented problem solving.
Studio
Model
Score
1
Claude Opus 5
100.0
2
GPT-5.6 Sol
93.0
3
Claude Fable 5
92.3
Agents
137 benchmarks
Covers tool use, web browsing, and long-horizon task execution.
Studio
Model
Score
1
Claude Fable 5
100.0
2
Kimi K3
97.1
3
GPT-5.6 Sol
96.7
Vision
108 benchmarks
Covers multimodal understanding, visual mathematics, screen, and video tasks.
Studio
Model
Score
1
Qwen3.8-Max
100.0
2
Claude Opus 5
95.5
3
Seed2.1 Pro
90.7
Long context
41 benchmarks
Covers retrieval, synthesis, and reasoning over million-token contexts.
Studio
Model
Score
1
Gemini 3.6 Flash
100.0
2
Muse Glimmer 30B
98.9
3
GPT-5.6 Sol
98.7
More domains
More data needed
Domains with limited samples do not show model rankings yet.
View benchmarks
Pending domains
Minimum 3
Writing
2 / 3
A domain joins the homepage ranking automatically after reaching the minimum coverage.