Reasoning benchmark

CritPt leaderboard

As of October 2026, GPT-5.6 Sol has the highest published CritPt score on Noometry at 32.3%, out of 134 models with results.

Last verified

About CritPt

Research-level physics problems written by working physicists.

Category
Reasoning
Introduced
2025
Format
Exact answer
Unit
Percent (random guessing ≈ 0%)
Official site
critpt.com

Top 15 models

Top models on CritPt
  1. GPT-5.6 Sol 32.3%
  2. Claude Opus 5.5 31.7%
  3. GPT-6.1 Sol 31.7%
  4. GPT-6 Astra 31.7%
  5. Claude Sonnet 5.5 31.4%
  6. Claude Fable 5.1 31.1%
  7. GPT-6 Sol 30.9%
  8. GPT-5.5 Pro 30.6%
  9. GPT-5.4 Pro 30%
  10. GPT-5.6 Terra 30%
  11. Claude Opus 5 29.1%
  12. Claude Fable 5 28.6%
  13. Gemini 4 Argon 27.1%
  14. GPT-5.5 27.1%
  15. MiMo-V2.6-Pro 26.6%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

CritPt results by model
#ModelProviderScoreSettingSourceDate
1GPT-5.6 Sol OpenAI32.3%maxEpoch AI
2Claude Opus 5.5 Anthropic31.7%maxEpoch AI
3GPT-6.1 Sol OpenAI31.7%maxEpoch AI
4GPT-6 Astra OpenAI31.7%maxEpoch AI
5Claude Sonnet 5.5 Anthropic31.4%maxEpoch AI
6Claude Fable 5.1 Anthropic31.1%xhighEpoch AI
7GPT-6 Sol OpenAI30.9%maxEpoch AI
8GPT-5.5 Pro OpenAI30.6%xhighEpoch AI
9GPT-5.4 Pro OpenAI30%xhighEpoch AI
10GPT-5.6 Terra OpenAI30%maxEpoch AI
11Claude Opus 5 Anthropic29.1%maxEpoch AI
12Claude Fable 5 Anthropic28.6%maxEpoch AI
13Gemini 4 Argon Google27.1%highEpoch AI
14GPT-5.5 OpenAI27.1%xhighEpoch AI
15MiMo-V2.6-Pro Xiaomi26.6%Epoch AI
16Muse Spark 1.3 Meta26%xhighEpoch AI
17Gemini 3 Deep Think Google25.7%Epoch AI
18GPT-5.4 OpenAI23.4%xhighEpoch AI
19Kimi K3 Moonshot AI23.4%maxEpoch AI
20Claude Opus 4.8 Anthropic20.9%maxEpoch AI
21GLM-5.2 Z.ai (Zhipu)20.9%maxEpoch AI
22Step 5 Preview StepFun20.9%Epoch AI
23GPT-5.6 Luna OpenAI20.6%maxEpoch AI
24Qwen3.8 Max Alibaba (Qwen)20%Epoch AI
25Grok 4.6 xAI19.7%xhighEpoch AI
26GPT-6 Luna OpenAI19.4%maxEpoch AI
27GLM-5.3 Z.ai (Zhipu)19.1%maxEpoch AI
28Gemini 3.8 Flash Google18.3%highEpoch AI
29DeepSeek V4 Pro DeepSeek18%maxEpoch AI
30Grok 4.7 xAI18%highEpoch AI
31Gemini 3.1 Pro Preview Google17.7%Epoch AI
32Muse Spark 1.2 Meta17.7%xhighEpoch AI
33Claude Sonnet 5 Anthropic16.9%maxEpoch AI
34DeepSeek V4 Flash DeepSeek16.6%maxEpoch AI
35GLM-5.3-Flash Z.ai (Zhipu)15.4%Epoch AI
36Grok 4.5 xAI15.4%highEpoch AI
37Muse Spark 1.1 Meta15.1%xhighEpoch AI
38DeepSeek V4.1 Flash DeepSeek14.3%maxEpoch AI
39Gemini 3.7 Flash Google14.3%highEpoch AI
40Qwen3.7 Max Alibaba (Qwen)13.4%Epoch AI
41Gemini 3.5 Flash Google13.1%highEpoch AI
42GPT-5 OpenAI12.6%highEpoch AI
43Claude Opus 4.7 Anthropic12%maxEpoch AI
44MiMo-V2.6-Flash Xiaomi12%Epoch AI
45Muse Spark Meta11.3%Epoch AI
46Gemini 3.6 Flash Google10.6%highEpoch AI
47GPT-5.4 mini OpenAI10%xhighEpoch AI
48Kimi K2.7 Code Moonshot AI10%Epoch AI
49GPT-5.4 nano OpenAI9.3%xhighEpoch AI
50Grok Build 0.1 xAI9.1%Epoch AI
51Qwen3.7 Plus Alibaba (Qwen)9.1%Epoch AI
52Inkling-Small Thinking Machines Lab8.3%Epoch AI
53Grok 4.3 xAI8%highEpoch AI
54Kimi K2.6 Moonshot AI8%Epoch AI
55Gemini 3 Pro Google6.9%Epoch AI
56Inkling Thinking Machines Lab5.4%xhighEpoch AI
57Qwen3.8 27B Alibaba (Qwen)5.4%xhighEpoch AI
58GPT-5.1 OpenAI4.9%Epoch AI
59GLM-5.1 Z.ai (Zhipu)4.6%Epoch AI
60MiMo-V2.5-Pro Xiaomi4%Epoch AI
61MiMo-V2.5 Xiaomi3.7%Epoch AI
62MiniMax-M3 MiniMax3.7%Epoch AI
63Ring-2.6-1T Ant Group (inclusionAI)3.7%Epoch AI
64Claude Sonnet 4.6 Anthropic3.1%maxEpoch AI
65Kimi K2.5 Moonshot AI3.1%Epoch AI
66Nemotron 3 Super NVIDIA3.1%Epoch AI
67Nemotron 3 Ultra NVIDIA3.1%Epoch AI
68DeepSeek-V3.2-Exp DeepSeek2.9%thinkingEpoch AI
69Qwen3.6 Plus Alibaba (Qwen)2.9%Epoch AI
70Muse Glimmer Meta2.6%highEpoch AI
71Step 3.7 Flash StepFun2.3%Epoch AI
72Gemini 2.5 Pro Google2%Epoch AI
73DeepSeek-V3.1-Terminus DeepSeek1.7%Epoch AI
74GLM-4.7 Z.ai (Zhipu)1.7%Epoch AI
75Gemma 4 31B IT Google1.4%Epoch AI
76gpt-oss-20b OpenAI1.4%highEpoch AI
77o3 OpenAI1.4%highEpoch AI
78Claude Sonnet 4.5 Anthropic1.1%Epoch AI
79Gemini 3.1 Flash Lite Google1.1%Epoch AI
80GLM-4.6 Z.ai (Zhipu)1.1%Epoch AI
81gpt-oss-120b OpenAI1.1%highEpoch AI
82DeepSeek-R1 DeepSeek1.1%Epoch AI
83Gemini 2.5 Flash Google1.1%Epoch AI
84Qwen3.5 122B-A10B Alibaba (Qwen)0.9%noneEpoch AI
85Qwen3.6 27B Alibaba (Qwen)0.9%noneEpoch AI
86Trinity Large Thinking Arcee AI0.9%Epoch AI
87Mercury 2 Inception0.8%Epoch AI
88o4-mini OpenAI0.6%highEpoch AI
89MiniMax-M2.7 MiniMax0.6%Epoch AI
90Qwen3.5 35B-A3B Alibaba (Qwen)0.6%noneEpoch AI
91Claude Opus 4 Anthropic0.3%Epoch AI
92Claude Sonnet 4 Anthropic0.3%Epoch AI
93Command A Plus Cohere0.3%Epoch AI
94Magistral Medium Mistral AI0.3%Epoch AI
95Magistral Small Mistral AI0.3%Epoch AI
96o3-mini OpenAI0.3%highEpoch AI
97Qwen3-30B-A3B Alibaba (Qwen)0.3%Epoch AI
98Qwen3 32B Alibaba (Qwen)0.3%Epoch AI
99Qwen3.5-9B Alibaba (Qwen)0.3%Epoch AI
100Qwen3.6 35B-A3B Alibaba (Qwen)0.3%Epoch AI
101Claude 3.5 Haiku Anthropic0%Epoch AI
102Claude Haiku 4.5 Anthropic0%Epoch AI
103DeepSeek-V3 DeepSeek0%Epoch AI
104DeepSeek-V3 DeepSeek0%Epoch AI
105Devstral Small 2505 Mistral AI0%Epoch AI
106Gemini 3.5 Flash Lite Google0%Epoch AI
107Gemma 3 12B Google0%Epoch AI
108Gemma 3 27B Google0%Epoch AI
109Gemma 4 26B A4B IT Google0%Epoch AI
110GPT-4.1 mini OpenAI0%Epoch AI
111GPT-4.1 nano OpenAI0%Epoch AI
112GPT-4o OpenAI0%Epoch AI
113GPT-5.5 Instant OpenAI0%Epoch AI
114GPT-5 Mini OpenAI0%Epoch AI
115Granite 4.1 30B IBM0%Epoch AI
116Llama 3.1-8B Meta0%Epoch AI
117Llama-3.3-70B-Instruct Meta0%Epoch AI
118Llama 4 Maverick Meta0%Epoch AI
119Llama 4 Maverick Meta0%Epoch AI
120Llama 4 Scout Meta0%Epoch AI
121Mercury 2.5 Inception0%Epoch AI
122MiMo-V2-Flash Xiaomi0%Epoch AI
123Mistral Large Mistral AI0%Epoch AI
124Mistral Medium Mistral AI0%Epoch AI
125Mistral Medium Mistral AI0%Epoch AI
126Mistral Small Mistral AI0%Epoch AI
127Mistral Small Mistral AI0%Epoch AI
128Nova 2.0 Pro Preview Amazon0%lowEpoch AI
129Phi-4 Mini Microsoft0%Epoch AI
130Qwen3 14B Alibaba (Qwen)0%Epoch AI
131Qwen3 235B-A22B Alibaba (Qwen)0%Epoch AI
132Qwen3 8B Alibaba (Qwen)0%Epoch AI
133Qwen3 Coder Next Alibaba (Qwen)0%Epoch AI
134Solar Pro 3 Upstage0%Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does CritPt measure?

Research-level physics problems written by working physicists.

Which model has the highest CritPt score?

As of October 2026, GPT-5.6 Sol has the highest published CritPt score on Noometry at 32.3%, out of 134 models with results.

What is the best open-weight model on CritPt?

MiMo-V2.6-Pro has the highest CritPt accuracy among open-weight models at 26.6%, ranking 15 of 134 overall.