Reasoning benchmark

# EBR-Bench leaderboard

> EBR-Bench results for 24 AI models, led by GPT-6 Astra at 76.2%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/ebr-bench
- Last updated: 2026-10-10
- Title: EBR-Bench Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-6 Astra has the highest published EBR-Bench score on Noometry at 76.2%, out of 24 models with results.

Last verified October 10, 2026

## About EBR-Bench

A description with primary sources is being prepared for this benchmark.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2026
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [epoch.ai](https://epoch.ai/benchmarks)

## Top 15 models

Top models on EBR-Bench

1.  GPT-6 Astra 76.2%
2.  Claude Opus 5.5 71.4%
3.  Claude Fable 5.1 57.1%
4.  GPT-6.1 Sol 54.3%
5.  GPT-6 Sol 53.3%
6.  Claude Opus 5 45.7%
7.  GPT-5.6 Sol 44.8%
8.  Claude Fable 5 39.5%
9.  GPT-5.5 34.3%
10.  Grok 4.6 30.5%
11.  Claude Opus 4.8 28.6%
12.  GPT-5.4 25.4%
13.  GPT-5.2 23%
14.  Claude Opus 4.7 19%
15.  Claude Opus 4.5 14.3%
16.  020406080

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

EBR-Bench results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | [OpenAI](https://noometry.com/providers/openai) | 76.2% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-30 |
| 2 | [Claude Opus 5.5](https://noometry.com/models/claude-opus-5-5) | [Anthropic](https://noometry.com/providers/anthropic) | 71.4% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-09-22 |
| 3 | [Claude Fable 5.1](https://noometry.com/models/claude-fable-5-1) | [Anthropic](https://noometry.com/providers/anthropic) | 57.1% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-09-08 |
| 4 | [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol) | [OpenAI](https://noometry.com/providers/openai) | 54.3% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-09-29 |
| 5 | [GPT-6 Sol](https://noometry.com/models/gpt-6-sol) | [OpenAI](https://noometry.com/providers/openai) | 53.3% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-09-22 |
| 6 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | [Anthropic](https://noometry.com/providers/anthropic) | 45.7% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-09-05 |
| 7 | [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol) | [OpenAI](https://noometry.com/providers/openai) | 44.8% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-09-04 |
| 8 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | [Anthropic](https://noometry.com/providers/anthropic) | 39.5% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-07-17 |
| 9 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | [OpenAI](https://noometry.com/providers/openai) | 34.3% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) | 2026-07-27 |
| 10 | [Grok 4.6](https://noometry.com/models/grok-4-6) | [xAI](https://noometry.com/providers/xai) | 30.5% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-18 |
| 11 | [Claude Opus 4.8](https://noometry.com/models/claude-opus-4-8) | [Anthropic](https://noometry.com/providers/anthropic) | 28.6% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-07 |
| 12 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | [OpenAI](https://noometry.com/providers/openai) | 25.4% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 13 | [GPT-5.2](https://noometry.com/models/gpt-5-2) | [OpenAI](https://noometry.com/providers/openai) | 23% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-26 |
| 14 | [Claude Opus 4.7](https://noometry.com/models/claude-opus-4-7) | [Anthropic](https://noometry.com/providers/anthropic) | 19% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-30 |
| 15 | [Claude Opus 4.5](https://noometry.com/models/claude-opus-4-5) | [Anthropic](https://noometry.com/providers/anthropic) | 14.3% | 128K | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 16 | [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 14.3% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 17 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 12.7% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-29 |
| 18 | [GPT-5](https://noometry.com/models/gpt-5) | [OpenAI](https://noometry.com/providers/openai) | 12.7% | high | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-29 |
| 19 | [GLM-5.2](https://noometry.com/models/glm-5-2) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 9.5% | max | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-29 |
| 20 | [Qwen3.7 Max](https://noometry.com/models/qwen3-7-max) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 9.5% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 21 | [Claude Opus 4.1](https://noometry.com/models/claude-opus-4-1) | [Anthropic](https://noometry.com/providers/anthropic) | 7.9% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 22 | [Gemini 3.5 Flash](https://noometry.com/models/gemini-3-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 4.8% | high | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 23 | [Claude Sonnet 4.5](https://noometry.com/models/claude-sonnet-4-5) | [Anthropic](https://noometry.com/providers/anthropic) | 2.4% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |
| 24 | [Kimi K2.6](https://noometry.com/models/kimi-k2-6) | [Moonshot AI](https://noometry.com/providers/moonshot) | 2.4% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-06-25 |

## Compare the leaders

-   [GPT-6 Astra vs Claude Opus 5.5](https://noometry.com/compare/claude-opus-5-5-vs-gpt-6-astra)
-   [GPT-6 Astra vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-gpt-6-astra)
-   [GPT-6 Astra vs GPT-6.1 Sol](https://noometry.com/compare/gpt-6-1-sol-vs-gpt-6-astra)
-   [GPT-6 Astra vs GPT-6 Sol](https://noometry.com/compare/gpt-6-astra-vs-gpt-6-sol)
-   [Claude Opus 5.5 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-claude-opus-5-5)
-   [Claude Opus 5.5 vs GPT-6.1 Sol](https://noometry.com/compare/claude-opus-5-5-vs-gpt-6-1-sol)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)
-   [Mystery Game Puzzles](https://noometry.com/benchmarks/mystery-game-puzzles)

## Frequently asked questions

### Which model has the highest EBR-Bench score?

As of October 2026, GPT-6 Astra has the highest published EBR-Bench score on Noometry at 76.2%, out of 24 models with results.

### What is the best open-weight model on EBR-Bench?

GLM-5.2 has the highest EBR-Bench accuracy among open-weight models at 9.5%, ranking 19 of 24 overall.

### Cite this page

Noometry. (2026). EBR-Bench leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/ebr-bench

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/ebr-bench.md).
