Agentic & Tool Use benchmark

METR Time Horizons leaderboard

As of October 2026, Claude Mythos Preview has the highest published METR Time Horizons score on Noometry at 85.2%, out of 32 models with results.

Last verified

About METR Time Horizons

METR's measure of how long a task (in human expert time) a model can complete with 50% reliability.

Category
Agentic & Tool Use
Introduced
2025
Format
Software tasks
Unit
Percent (random guessing ≈ 0%)
Official site
metr.org

Top 15 models

Top models on METR Time Horizons
  1. Claude Mythos Preview 85.2%
  2. Claude Opus 4.6 78.9%
  3. Gemini 3.1 Pro Preview 77%
  4. GPT-5.2 75.3%
  5. Claude Opus 4.5 75%
  6. GPT-5.3 Codex 74.5%
  7. GPT-5.4 74.3%
  8. Gemini 3 Pro 71%
  9. GPT-5.1-Codex 70.8%
  10. GPT-5 69.6%
  11. Claude Sonnet 4.5 67.4%
  12. Claude Opus 4.1 66.8%
  13. Grok 4 66.6%
  14. o3 65.4%
  15. Claude Opus 4 63.9%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does METR Time Horizons measure?

METR's measure of how long a task (in human expert time) a model can complete with 50% reliability.

Which model has the highest METR Time Horizons score?

As of October 2026, Claude Mythos Preview has the highest published METR Time Horizons score on Noometry at 85.2%, out of 32 models with results.

What is the best open-weight model on METR Time Horizons?

Kimi K2 (Jul 2025) has the highest METR Time Horizons accuracy among open-weight models at 59.2%, ranking 19 of 32 overall.