Long Context benchmark

# CL-bench leaderboard

> CL-bench results for 19 AI models, led by GPT-5.4 at 27.9%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/cl-bench
- Last updated: 2026-10-10
- Title: CL-bench Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-5.4 has the highest published CL-bench score on Noometry at 27.9%, out of 19 models with results.

Last verified October 10, 2026

## About CL-bench

A description with primary sources is being prepared for this benchmark.

- **Category:** [Long Context](https://noometry.com/best/long-context)
- **Introduced:** 2026
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [epoch.ai](https://epoch.ai/benchmarks)

## Top 15 models

Top models on CL-bench

1.  GPT-5.4 27.9%
2.  GPT-5.1 23.7%
3.  Grok 4.20 (Non-Reasoning) 22.2%
4.  Claude Opus 4.5 21.1%
5.  Gemini 3.1 Pro Preview 20.8%
6.  Claude Opus 4.6 20.7%
7.  Qwen3.6 Plus 20.3%
8.  Qwen3.5 Plus 19.8%
9.  Kimi K2.5 19.3%
10.  GLM-5 18.7%
11.  GPT-5.2 18.2%
12.  o3 17.8%
13.  Kimi K2 (Jul 2025) 17.6%
14.  GLM-4.7 15.9%
15.  Gemini 3 Pro 15.8%
16.  1015202530

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

CL-bench results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | [OpenAI](https://noometry.com/providers/openai) | 27.9% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | [OpenAI](https://noometry.com/providers/openai) | 23.7% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20) | [xAI](https://noometry.com/providers/xai) | 22.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Claude Opus 4.5](https://noometry.com/models/claude-opus-4-5) | [Anthropic](https://noometry.com/providers/anthropic) | 21.1% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 20.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 20.7% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [Qwen3.6 Plus](https://noometry.com/models/qwen3-6-plus) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 20.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [Qwen3.5 Plus](https://noometry.com/models/qwen3-5-plus) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 19.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Kimi K2.5](https://noometry.com/models/kimi-k2-5) | [Moonshot AI](https://noometry.com/providers/moonshot) | 19.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 10 | [GLM-5](https://noometry.com/models/glm-5) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 18.7% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 11 | [GPT-5.2](https://noometry.com/models/gpt-5-2) | [OpenAI](https://noometry.com/providers/openai) | 18.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 12 | [o3](https://noometry.com/models/o3) | [OpenAI](https://noometry.com/providers/openai) | 17.8% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 13 | [Kimi K2 (Jul 2025)](https://noometry.com/models/kimi-k2) | [Moonshot AI](https://noometry.com/providers/moonshot) | 17.6% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 14 | [GLM-4.7](https://noometry.com/models/glm-4-7) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 15.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 15 | [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 15.8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 16 | [MiMo-V2-Pro](https://noometry.com/models/mimo-v2-pro) | [Xiaomi](https://noometry.com/providers/xiaomi) | 15.7% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 17 | [Qwen3 Max](https://noometry.com/models/qwen3-max) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 14.5% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 18 | [DeepSeek-V3.2-Exp](https://noometry.com/models/deepseek-v3-2-exp) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 13.2% | thinking | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 19 | [MiniMax-M2.5](https://noometry.com/models/minimax-m2-5) |  [![](/logos/minimax.svg) MiniMax](https://noometry.com/providers/minimax) | 11.4% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [GPT-5.4 vs GPT-5.1](https://noometry.com/compare/gpt-5-1-vs-gpt-5-4)
-   [GPT-5.4 vs Grok 4.20 (Non-Reasoning)](https://noometry.com/compare/gpt-5-4-vs-grok-4-20)
-   [GPT-5.4 vs Claude Opus 4.5](https://noometry.com/compare/claude-opus-4-5-vs-gpt-5-4)
-   [GPT-5.4 vs Gemini 3.1 Pro Preview](https://noometry.com/compare/gemini-3-1-pro-preview-vs-gpt-5-4)
-   [GPT-5.1 vs Grok 4.20 (Non-Reasoning)](https://noometry.com/compare/gpt-5-1-vs-grok-4-20)
-   [GPT-5.1 vs Claude Opus 4.5](https://noometry.com/compare/claude-opus-4-5-vs-gpt-5-1)

## Other long context benchmarks

-   [Fiction.LiveBench](https://noometry.com/benchmarks/fiction-livebench)
-   [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query)
-   [CL-bench Life](https://noometry.com/benchmarks/cl-bench-life)

## Frequently asked questions

### Which model has the highest CL-bench score?

As of October 2026, GPT-5.4 has the highest published CL-bench score on Noometry at 27.9%, out of 19 models with results.

### What is the best open-weight model on CL-bench?

Kimi K2.5 has the highest CL-bench accuracy among open-weight models at 19.3%, ranking 9 of 19 overall.

### Cite this page

Noometry. (2026). CL-bench leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/cl-bench

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/cl-bench.md).
