Long Context benchmark

CL-bench leaderboard

As of October 2026, GPT-5.4 has the highest published CL-bench score on Noometry at 27.9%, out of 19 models with results.

Last verified

About CL-bench

A description with primary sources is being prepared for this benchmark.

Category
Long Context
Introduced
2026
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 15 models

Top models on CL-bench
  1. GPT-5.4 27.9%
  2. GPT-5.1 23.7%
  3. Grok 4.20 (Non-Reasoning) 22.2%
  4. Claude Opus 4.5 21.1%
  5. Gemini 3.1 Pro Preview 20.8%
  6. Claude Opus 4.6 20.7%
  7. Qwen3.6 Plus 20.3%
  8. Qwen3.5 Plus 19.8%
  9. Kimi K2.5 19.3%
  10. GLM-5 18.7%
  11. GPT-5.2 18.2%
  12. o3 17.8%
  13. Kimi K2 (Jul 2025) 17.6%
  14. GLM-4.7 15.9%
  15. Gemini 3 Pro 15.8%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other long context benchmarks

Frequently asked questions

Which model has the highest CL-bench score?

As of October 2026, GPT-5.4 has the highest published CL-bench score on Noometry at 27.9%, out of 19 models with results.

What is the best open-weight model on CL-bench?

Kimi K2.5 has the highest CL-bench accuracy among open-weight models at 19.3%, ranking 9 of 19 overall.