Agentic & Tool Use benchmark
OSWorld 2.0 leaderboard
As of October 2026, Claude Opus 5 has the highest published OSWorld 2.0 score on Noometry at 31.4%, out of 9 models with results.
Last verified
About OSWorld 2.0
The second version of the OSWorld computer-use benchmark.
- Category
- Agentic & Tool Use
- Introduced
- 2026
- Format
- Computer use
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- os-world.github.io
Top 9 models
- Claude Opus 5 31.4%
- GPT-5.6 Sol 27.3%
- Claude Opus 4.8 20.6%
- Claude Opus 4.7 18.2%
- GPT-5.5 13%
- Claude Sonnet 4.6 9.3%
- Kimi K2.6 4.6%
- MiniMax-M3 4.6%
- Qwen3.7 Plus 2.8%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 31.4% | max | Epoch AI | |
| 2 | GPT-5.6 Sol | OpenAI | 27.3% | max | Epoch AI | |
| 3 | Claude Opus 4.8 | Anthropic | 20.6% | max | Epoch AI | |
| 4 | Claude Opus 4.7 | Anthropic | 18.2% | max | Epoch AI | |
| 5 | GPT-5.5 | OpenAI | 13% | xhigh | Epoch AI | |
| 6 | Claude Sonnet 4.6 | Anthropic | 9.3% | medium | Epoch AI | |
| 7 | Kimi K2.6 | Moonshot AI | 4.6% | Epoch AI | ||
| 8 | MiniMax-M3 | 4.6% | Epoch AI | |||
| 9 | Qwen3.7 Plus | 2.8% | Epoch AI |
Compare the leaders
Frequently asked questions
What does OSWorld 2.0 measure?
The second version of the OSWorld computer-use benchmark.
Which model has the highest OSWorld 2.0 score?
As of October 2026, Claude Opus 5 has the highest published OSWorld 2.0 score on Noometry at 31.4%, out of 9 models with results.
What is the best open-weight model on OSWorld 2.0?
Kimi K2.6 has the highest OSWorld 2.0 accuracy among open-weight models at 4.6%, ranking 7 of 9 overall.