Skip to main content
leaderboard
.cn
Toggle Sidebar
Search models or organizations…
⌘K
Explore
Frontier Models
Historical Trends
Local Models
Compare Models
Ranking Studio
Library
Models
Pricing
Organizations
Benchmarks
Data sources
Settings
Toggle Sidebar
leaderboard
.cn
Loading page content
Feedback
Benchmark details loading
Back to benchmarks
OSWorld 2.0 (Meta, Sep 2026)
Agents
Share
The source chart does not identify the task release, scoring mode, or agent harness. This result is kept separate from first-attempt, partial, and strict OSWorld 2.0 variants.
Ranking
Sources
#
Model
Score
1
Claude Opus 5
Anthropic · max mode
68.3
%
2
Muse Spark 1.3
Meta · max mode
66.9
%
3
GPT-5.6 Sol
OpenAI · max mode
62.7
%
4
Muse Spark 1.2
Meta · xhigh mode
47.6
%
4 models total