Reasoning benchmark

ARC-AGI-2 leaderboard

As of October 2026, GPT-6 Astra has the highest published ARC-AGI-2 score on Noometry at 95%, out of 83 models with results.

Last verified

About ARC-AGI-2

A harder successor to ARC-AGI with puzzles designed to resist brute-force search.

Category
Reasoning
Introduced
2025
Format
Grid puzzles
Unit
Percent (random guessing ≈ 0%)
Official site
arcprize.org

Top 15 models

Top models on ARC-AGI-2
  1. GPT-6 Astra 95%
  2. GPT-6.1 Sol 94.2%
  3. Claude Opus 5.5 93.3%
  4. GPT-5.6 Sol 92.5%
  5. Claude Opus 5 90.4%
  6. Claude Fable 5.1 90%
  7. GPT-6 Sol 89.6%
  8. Claude Fable 5 89.2%
  9. Gemini 3.8 Flash 89.2%
  10. GPT-5.5 85%
  11. Gemini 3.7 Flash 84.6%
  12. Gemini 3 Deep Think 84.6%
  13. GPT-5.5 Pro 84.6%
  14. GPT-5.6 Terra 83.9%
  15. GPT-5.4 Pro 83.3%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

ARC-AGI-2 results by model
#ModelProviderScoreSettingSourceDate
1GPT-6 Astra OpenAI95%maxEpoch AI
2GPT-6.1 Sol OpenAI94.2%maxEpoch AI
3Claude Opus 5.5 Anthropic93.3%highEpoch AI
4GPT-5.6 Sol OpenAI92.5%maxEpoch AI
5Claude Opus 5 Anthropic90.4%maxEpoch AI
6Claude Fable 5.1 Anthropic90%maxEpoch AI
7GPT-6 Sol OpenAI89.6%maxEpoch AI
8Claude Fable 5 Anthropic89.2%maxEpoch AI
9Gemini 3.8 Flash Google89.2%highEpoch AI
10GPT-5.5 OpenAI85%xhighEpoch AI
11Gemini 3.7 Flash Google84.6%highEpoch AI
12Gemini 3 Deep Think Google84.6%Epoch AI
13GPT-5.5 Pro OpenAI84.6%highEpoch AI
14GPT-5.6 Terra OpenAI83.9%maxEpoch AI
15GPT-5.4 Pro OpenAI83.3%xhighEpoch AI
16Gemini 3.1 Pro Preview Google77.1%Epoch AI
17Claude Opus 4.7 Anthropic75.8%maxEpoch AI
18GPT-5.4 OpenAI74%xhighEpoch AI
19Claude Opus 4.8 Anthropic72.1%highEpoch AI
20Gemini 3.5 Flash Google72.1%highEpoch AI
21Claude Opus 4.6 Anthropic69.2%120KEpoch AI
22Grok 4.6 xAI67.1%xhighEpoch AI
23GLM-5.3-Flash Z.ai (Zhipu)65.8%maxEpoch AI
24Grok 4.20 (Non-Reasoning) xAI65.1%Epoch AI
25DeepSeek V4 Flash DeepSeek61.4%maxEpoch AI
26DeepSeek V4 Pro DeepSeek61.3%maxEpoch AI
27Claude Sonnet 4.6 Anthropic60.4%highEpoch AI
28Gemini 3.6 Flash Google60.4%highEpoch AI
29Kimi K3 Moonshot AI60.4%maxEpoch AI
30GPT-5.6 Luna OpenAI59.5%maxEpoch AI
31GPT-6 Luna OpenAI59.3%maxEpoch AI
32GPT-5.2 Pro OpenAI54.2%highEpoch AI
33GPT-5.2 OpenAI52.9%xhighEpoch AI
34Grok 4.5 xAI52.6%highEpoch AI
35Qwen3.8 27B Alibaba (Qwen)42.4%xhighEpoch AI
36Inkling-Small Thinking Machines Lab40.1%xhighEpoch AI
37Claude Opus 4.5 Anthropic37.6%64KEpoch AI
38Inkling Thinking Machines Lab36.5%Epoch AI
39Gemini 3 Flash Preview Google33.6%Epoch AI
40Gemini 3 Pro Google31.1%Epoch AI
41GLM-5.2 Z.ai (Zhipu)22.8%Epoch AI
42GPT-5.4 mini OpenAI18.9%xhighEpoch AI
43GPT-5 Pro OpenAI18.3%Epoch AI
44GPT-5.1 OpenAI17.6%highEpoch AI
45Grok 4 xAI16%Epoch AI
46Claude Sonnet 4.5 Anthropic13.6%32KEpoch AI
47Kimi K2.5 Moonshot AI11.8%Epoch AI
48Gemini 3.5 Flash Lite Google10.3%highEpoch AI
49GPT-5 OpenAI9.9%highEpoch AI
50Claude Opus 4 Anthropic8.6%16KEpoch AI
51o3 OpenAI6.5%highEpoch AI
52o4-mini OpenAI6.1%highEpoch AI
53Claude Sonnet 4 Anthropic5.9%16KEpoch AI
54GPT-5.4 nano OpenAI5.7%xhighEpoch AI
55Grok 4 Fast xAI5.3%Epoch AI
56Gemini 2.5 Pro Google4.9%32KEpoch AI
57GLM-5 Z.ai (Zhipu)4.9%Epoch AI
58MiniMax-M2.5 MiniMax4.9%Epoch AI
59o3-pro OpenAI4.9%highEpoch AI
60GPT-5 Mini OpenAI4.4%highEpoch AI
61Claude Haiku 4.5 Anthropic4%32KEpoch AI
62DeepSeek-V3.2-Exp DeepSeek4%Epoch AI
63o3-mini OpenAI3%highEpoch AI
64GPT-5 Nano OpenAI2.6%highEpoch AI
65Gemini 2.5 Flash Google2.5%23KEpoch AI
66DeepSeek-R1 DeepSeek1.3%Epoch AI
67Gemini 2.0 Flash (Feb 2025) Google1.3%Epoch AI
68Qwen3 235B-A22B Alibaba (Qwen)1.3%Epoch AI
69Claude 3.7 Sonnet Anthropic0.9%8KEpoch AI
70o1-mini OpenAI0.8%Epoch AI
71Gemini 1.5 Pro (May 2024) Google0.8%Epoch AI
72GPT-4.5 OpenAI0.8%Epoch AI
73GPT-4.1 OpenAI0.4%Epoch AI
74Grok-3 mini xAI0.4%lowEpoch AI
75GPT-4.1 mini OpenAI0%Epoch AI
76GPT-4.1 nano OpenAI0%Epoch AI
77GPT-4o OpenAI0%Epoch AI
78GPT-4o mini OpenAI0%Epoch AI
79Grok 3 xAI0%Epoch AI
80Llama 4 Maverick Meta0%Epoch AI
81Llama 4 Scout Meta0%Epoch AI
82Magistral Medium Mistral AI0%Epoch AI
83Magistral Small Mistral AI0%Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does ARC-AGI-2 measure?

A harder successor to ARC-AGI with puzzles designed to resist brute-force search.

Which model has the highest ARC-AGI-2 score?

As of October 2026, GPT-6 Astra has the highest published ARC-AGI-2 score on Noometry at 95%, out of 83 models with results.

What is the best open-weight model on ARC-AGI-2?

GLM-5.3-Flash has the highest ARC-AGI-2 accuracy among open-weight models at 65.8%, ranking 23 of 83 overall.