Math benchmark

FrontierMath Tier 4 (v1) leaderboard

As of October 2026, AI Co-Mathematician has the highest published FrontierMath Tier 4 (v1) score on Noometry at 47.9%, out of 55 models with results.

Last verified

About FrontierMath Tier 4 (v1)

Earlier Tier 4 problem set, kept for historical comparison.

Category
Math
Introduced
2025
Format
Exact answer
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 15 models

Top models on FrontierMath Tier 4 (v1)
  1. AI Co-Mathematician 47.9%
  2. GPT-5.5 Pro 39.6%
  3. GPT-5.4 Pro 37.5%
  4. GPT-5.5 35.4%
  5. GPT-5.2 Pro 31.3%
  6. Claude Opus 4.8 31.3%
  7. GPT-5.4 27.1%
  8. Claude Opus 4.7 22.9%
  9. Claude Opus 4.6 22.9%
  10. GPT-5.2 18.8%
  11. Gemini 3 Pro 18.8%
  12. Gemini 3.1 Pro Preview 16.7%
  13. GPT-5 Pro 14.6%
  14. Muse Spark 14.6%
  15. Gemini 3.5 Flash 14.6%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

FrontierMath Tier 4 (v1) results by model
#ModelProviderScoreSettingSourceDate
1AI Co-Mathematician Google47.9%Epoch AI2026-05-08
2GPT-5.5 Pro OpenAI39.6%highEpoch AI2026-04-23
3GPT-5.4 Pro OpenAI37.5%Epoch AI2026-03-06
4GPT-5.5 OpenAI35.4%xhighEpoch AI2026-04-23
5GPT-5.2 Pro OpenAI31.3%Epoch AI2025-12-24
6Claude Opus 4.8 Anthropic31.3%maxEpoch AI2026-06-08
7GPT-5.4 OpenAI27.1%xhighEpoch AI2026-03-06
8Claude Opus 4.7 Anthropic22.9%xhighEpoch AI2026-04-17
9Claude Opus 4.6 Anthropic22.9%maxEpoch AI2026-02-12
10GPT-5.2 OpenAI18.8%highEpoch AI2025-12-11
11Gemini 3 Pro Google18.8%Epoch AI2025-11-21
12Gemini 3.1 Pro Preview Google16.7%Epoch AI2026-02-19
13GPT-5 Pro OpenAI14.6%highEpoch AI2025-10-09
14Muse Spark Meta14.6%Epoch AI2026-04-08
15Gemini 3.5 Flash Google14.6%highEpoch AI2026-05-25
16Kimi K2.6 Moonshot AI14.6%Epoch AI2026-05-08
17GLM-5.1 Z.ai (Zhipu)12.5%Epoch AI2026-05-12
18GPT-5 OpenAI12.5%highEpoch AI2025-10-30
19GPT-5.1 OpenAI12.5%highEpoch AI2025-11-13
20Gemini 2.5 Deep Think Google10.4%Epoch AI
21Qwen3.6 Plus Alibaba (Qwen)8.3%Epoch AI2026-05-12
22Claude Sonnet 4.6 Anthropic8.3%16KEpoch AI2026-02-21
23GPT-5.4 nano OpenAI6.3%highEpoch AI2026-04-15
24GPT-5 Mini OpenAI6.3%highEpoch AI2025-10-30
25o4-mini OpenAI6.3%highEpoch AI2025-07-01
26Kimi K2.5 Moonshot AI4.2%Epoch AI2026-02-02
27Claude Opus 4 Anthropic4.2%27KEpoch AI2025-07-01
28Claude Opus 4.1 Anthropic4.2%27KEpoch AI2025-08-05
29Claude Opus 4.5 Anthropic4.2%Epoch AI2025-11-25
30Claude Sonnet 4.5 Anthropic4.2%32KEpoch AI2025-10-22
31Gemini 2.5 Flash Google4.2%Epoch AI2025-12-18
32Gemini 2.5 Pro Google4.2%Epoch AI2025-07-03
33Gemini 3 Flash Preview Google4.2%Epoch AI2025-12-17
34o3-mini OpenAI4.2%highEpoch AI2025-07-01
35Qwen3.6 Max Preview Alibaba (Qwen)4.2%Epoch AI2026-05-28
36GLM-4.6 Z.ai (Zhipu)2.1%Epoch AI2025-12-08
37DeepSeek-V3.2-Exp DeepSeek2.1%Epoch AI2025-12-16
38GLM-5 Z.ai (Zhipu)2.1%Epoch AI2026-02-19
39Claude Haiku 4.5 Anthropic2.1%32KEpoch AI2025-10-22
40GPT-5 Nano OpenAI2.1%mediumEpoch AI2025-08-07
41Grok 4 xAI2.1%Epoch AI2025-08-11
42o3 OpenAI2.1%highEpoch AI2025-07-01
43Qwen3.5 Plus Alibaba (Qwen)2.1%Epoch AI2026-05-15
44GPT-5.4 mini OpenAI2.1%highEpoch AI2026-04-15
45Grok 4 Heavy xAI2.1%Epoch AI2025-10-09
46Claude 3.5 Sonnet Anthropic0%Epoch AI2025-07-01
47Claude 3.5 Sonnet Anthropic0%Epoch AI2025-07-01
48Claude Sonnet 4 Anthropic0%Epoch AI2025-07-01
49GLM-4.7 Z.ai (Zhipu)0%Epoch AI2026-01-30
50GPT-4.1 OpenAI0%Epoch AI2025-07-01
51Grok 3 xAI0%Epoch AI2025-07-01
52Kimi K2 (Jul 2025) Moonshot AI0%Epoch AI2025-12-04
53Qwen3 235B-A22B Alibaba (Qwen)0%Epoch AI2025-12-11
54Qwen3.5-Flash Alibaba (Qwen)0%Epoch AI2026-05-12
55Qwen3.6 Flash Alibaba (Qwen)0%Epoch AI2026-05-12

Compare the leaders

Other math benchmarks

Frequently asked questions

What does FrontierMath Tier 4 (v1) measure?

Earlier Tier 4 problem set, kept for historical comparison.

Which model has the highest FrontierMath Tier 4 (v1) score?

As of October 2026, AI Co-Mathematician has the highest published FrontierMath Tier 4 (v1) score on Noometry at 47.9%, out of 55 models with results.

What is the best open-weight model on FrontierMath Tier 4 (v1)?

Kimi K2.6 has the highest FrontierMath Tier 4 (v1) accuracy among open-weight models at 14.6%, ranking 16 of 55 overall.