Math benchmark

FrontierMath (Feb 2025 set) leaderboard

As of October 2026, GPT-5.5 Pro has the highest published FrontierMath (Feb 2025 set) score on Noometry at 52.4%, out of 68 models with results.

Last verified

About FrontierMath (Feb 2025 set)

Earlier FrontierMath problem set, kept for historical comparison.

Category
Math
Introduced
2025
Format
Exact answer
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 15 models

Top models on FrontierMath (Feb 2025 set)
  1. GPT-5.5 Pro 52.4%
  2. GPT-5.5 51.7%
  3. GPT-5.4 Pro 50%
  4. GPT-5.4 47.6%
  5. Claude Opus 4.8 47.2%
  6. Claude Opus 4.7 43.8%
  7. Claude Opus 4.6 40.7%
  8. GPT-5.2 40.7%
  9. Muse Spark 39%
  10. Gemini 3.5 Flash 39%
  11. Kimi K2.6 39%
  12. Gemini 3 Pro 37.6%
  13. Gemini 3.1 Pro Preview 36.9%
  14. Gemini 3 Flash Preview 35.6%
  15. GLM-5.1 33.4%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

FrontierMath (Feb 2025 set) results by model
#ModelProviderScoreSettingSourceDate
1GPT-5.5 Pro OpenAI52.4%highEpoch AI2026-04-23
2GPT-5.5 OpenAI51.7%xhighEpoch AI2026-04-23
3GPT-5.4 Pro OpenAI50%xhighEpoch AI2026-03-06
4GPT-5.4 OpenAI47.6%xhighEpoch AI2026-03-06
5Claude Opus 4.8 Anthropic47.2%maxEpoch AI2026-06-08
6Claude Opus 4.7 Anthropic43.8%xhighEpoch AI2026-04-17
7Claude Opus 4.6 Anthropic40.7%maxEpoch AI2026-02-12
8GPT-5.2 OpenAI40.7%xhighEpoch AI2025-12-13
9Muse Spark Meta39%Epoch AI2026-04-08
10Gemini 3.5 Flash Google39%highEpoch AI2026-05-22
11Kimi K2.6 Moonshot AI39%Epoch AI2026-05-07
12Gemini 3 Pro Google37.6%Epoch AI2025-11-21
13Gemini 3.1 Pro Preview Google36.9%Epoch AI2026-02-19
14Gemini 3 Flash Preview Google35.6%Epoch AI2025-12-17
15GLM-5.1 Z.ai (Zhipu)33.4%Epoch AI2026-05-11
16GPT-5 OpenAI32.4%highEpoch AI2025-11-13
17Claude Sonnet 4.6 Anthropic32.4%16KEpoch AI2026-02-20
18GPT-5.1 OpenAI31%highEpoch AI2025-11-13
19Gemini 2.5 Deep Think Google29%Epoch AI
20GPT-5.4 mini OpenAI28.3%highEpoch AI2026-04-15
21Kimi K2.5 Moonshot AI27.9%Epoch AI2026-02-03
22GPT-5 Mini OpenAI27.2%highEpoch AI2025-11-13
23Qwen3.6 Plus Alibaba (Qwen)26.2%Epoch AI2026-05-12
24GPT-5.4 nano OpenAI25.9%highEpoch AI2026-04-15
25o4-mini OpenAI24.8%highEpoch AI2025-11-13
26Qwen3.6 Max Preview Alibaba (Qwen)23.1%Epoch AI2026-05-27
27DeepSeek-V3.2-Exp DeepSeek22.1%Epoch AI2025-12-22
28Kimi K2 (Jul 2025) Moonshot AI21.4%Epoch AI2025-12-05
29Qwen3.5 Plus Alibaba (Qwen)21%Epoch AI2026-05-14
30Claude Opus 4.5 Anthropic20.7%Epoch AI2025-11-25
31Grok 4 xAI19.7%Epoch AI2025-11-13
32o3 OpenAI18.7%highEpoch AI2025-11-16
33GLM-5 Z.ai (Zhipu)16.4%Epoch AI2026-02-19
34Claude Sonnet 4.5 Anthropic15.2%32KEpoch AI2025-11-16
35Gemini 2.5 Pro Google14.1%Epoch AI2025-11-24
36o3-mini OpenAI12.4%highEpoch AI2025-11-16
37Qwen3.6 Flash Alibaba (Qwen)10.3%Epoch AI2026-05-12
38o1 OpenAI9.3%highEpoch AI2025-03-07
39Qwen3 235B-A22B Alibaba (Qwen)8.5%Epoch AI2025-12-11
40GPT-5 Nano OpenAI8.3%highEpoch AI2025-10-30
41Claude Opus 4.1 Anthropic7.2%27KEpoch AI2025-08-05
42Qwen3.5-Flash Alibaba (Qwen)6.2%Epoch AI2026-05-12
43Claude Haiku 4.5 Anthropic5.9%32KEpoch AI2025-10-22
44Grok-3 mini xAI5.9%highEpoch AI2025-04-10
45GPT-4.1 OpenAI5.5%Epoch AI2025-04-14
46Gemini 2.5 Flash Google4.8%Epoch AI2025-12-18
47Claude Opus 4 Anthropic4.5%Epoch AI2025-07-04
48GPT-4.1 mini OpenAI4.5%Epoch AI2025-04-14
49Claude 3.7 Sonnet Anthropic4.1%16KEpoch AI2025-03-06
50Claude Sonnet 4 Anthropic4.1%Epoch AI2025-07-04
51GLM-4.6 Z.ai (Zhipu)3.8%Epoch AI2025-12-08
52Grok 3 xAI3.8%Epoch AI2025-04-10
53GLM-4.7 Z.ai (Zhipu)2.4%Epoch AI2026-01-30
54Claude 3.5 Sonnet Anthropic2.1%Epoch AI2025-03-06
55DeepSeek-V3 DeepSeek1.7%Epoch AI2025-03-07
56Gemini 2.0 Flash (Feb 2025) Google1.7%Epoch AI2025-03-09
57o1-mini OpenAI1.7%mediumEpoch AI2025-03-06
58Qwen Plus Alibaba (Qwen)1.7%Epoch AI2025-05-14
59GPT-4.1 nano OpenAI1%Epoch AI2025-04-14
60Qwen Max Alibaba (Qwen)1%Epoch AI2025-04-02
61Grok-2 (Dec 2024) xAI0.7%Epoch AI2025-03-06
62Llama 4 Maverick Meta0.7%Epoch AI2025-04-08
63Mistral Medium Mistral AI0.3%Epoch AI2025-05-07
64Claude 3.5 Haiku Anthropic0.3%Epoch AI2025-03-07
65GPT-4o OpenAI0.3%Epoch AI2025-03-07
66Mistral Large Mistral AI0.3%Epoch AI2025-03-06
67Gemini 1.5 Flash (May 2024) Google0%Epoch AI2025-03-07
68Llama 4 Scout Meta0%Epoch AI2025-04-08

Compare the leaders

Other math benchmarks

Frequently asked questions

What does FrontierMath (Feb 2025 set) measure?

Earlier FrontierMath problem set, kept for historical comparison.

Which model has the highest FrontierMath (Feb 2025 set) score?

As of October 2026, GPT-5.5 Pro has the highest published FrontierMath (Feb 2025 set) score on Noometry at 52.4%, out of 68 models with results.

What is the best open-weight model on FrontierMath (Feb 2025 set)?

Kimi K2.6 has the highest FrontierMath (Feb 2025 set) accuracy among open-weight models at 39%, ranking 11 of 68 overall.