Agentic & Tool Use benchmark
METR Time Horizons leaderboard
As of October 2026, Claude Mythos Preview has the highest published METR Time Horizons score on Noometry at 85.2%, out of 32 models with results.
Last verified
About METR Time Horizons
METR's measure of how long a task (in human expert time) a model can complete with 50% reliability.
- Category
- Agentic & Tool Use
- Introduced
- 2025
- Format
- Software tasks
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- metr.org
Top 15 models
- Claude Mythos Preview 85.2%
- Claude Opus 4.6 78.9%
- Gemini 3.1 Pro Preview 77%
- GPT-5.2 75.3%
- Claude Opus 4.5 75%
- GPT-5.3 Codex 74.5%
- GPT-5.4 74.3%
- Gemini 3 Pro 71%
- GPT-5.1-Codex 70.8%
- GPT-5 69.6%
- Claude Sonnet 4.5 67.4%
- Claude Opus 4.1 66.8%
- Grok 4 66.6%
- o3 65.4%
- Claude Opus 4 63.9%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does METR Time Horizons measure?
METR's measure of how long a task (in human expert time) a model can complete with 50% reliability.
Which model has the highest METR Time Horizons score?
As of October 2026, Claude Mythos Preview has the highest published METR Time Horizons score on Noometry at 85.2%, out of 32 models with results.
What is the best open-weight model on METR Time Horizons?
Kimi K2 (Jul 2025) has the highest METR Time Horizons accuracy among open-weight models at 59.2%, ranking 19 of 32 overall.