Skip to main content
leaderboard
.cn
Toggle Sidebar
Search models or organizations…
⌘K
Explore
Frontier Models
Historical Trends
Local Models
Compare Models
Ranking Studio
Library
Models
Pricing
Organizations
Benchmarks
Data sources
Settings
Toggle Sidebar
leaderboard
.cn
Loading page content
Feedback
Loading page content
Back to models
o4-mini
Closed
Share
Add to comparison
Tool use · Reasoning
Overview
Evaluations
Evaluation overview
Overall rank
#89
67.8 points · 180 models
Coding rank
#64
72.3 points · 160 models
Data coverage
29 items
2 sources · updated through 2025-08-07
Domain performance
Coding
#64 / 160
72.3
Research
#78 / 174
79
Agents
#66 / 125
67.6
Vision
#48 / 65
72.7
Long context
#66 / 82
63.6
Evaluation highlights
All evaluations
AIME 2025
Reasoning · OpenAI GPT-5 Developer Release
92.7%
10
/ 60
TAU2-Airline
Agents · OpenAI GPT-5 Developer Release
60.2%
4
/ 17
Aider Polyglot
Coding · OpenAI GPT-5 Developer Release
58.2%
8
/ 29
OpenAI MRCR 2-needle 128k
Long context · OpenAI GPT-5 Developer Release
56.4%
4
/ 13
VideoMME long (with subtitles)
Multimodal · OpenAI GPT-5 Developer Release
79.5%
3
/ 8
COLLIE
Instruction following · OpenAI GPT-5 Developer Release
96.1%
6
/ 13
Basic information
Context (input / output)
200K / 100K
Release date
2025-04
Input modalities
Text · Image
Output modalities
Text
Related links
Official site