Coding benchmark

WeirdML leaderboard

As of October 2026, GPT-6 Astra has the highest published WeirdML score on Noometry at 93.6%, out of 119 models with results.

Last verified

About WeirdML

Unusual machine-learning tasks the model must solve by writing and iterating on working PyTorch code.

Category
Coding
Introduced
2025
Format
Code + execution
Unit
Percent (random guessing ≈ 0%)
Official site
htihle.github.io

Top 15 models

Top models on WeirdML
  1. GPT-6 Astra 93.6%
  2. Claude Fable 5.1 92.9%
  3. Claude Fable 5 91.9%
  4. Claude Opus 5 91.8%
  5. GPT-5.6 Sol 89.4%
  6. GPT-5.5 84.9%
  7. Gemini 3.8 Flash 84.8%
  8. Claude Opus 4.8 82.9%
  9. Kimi K3 82.6%
  10. GPT-5.3 Codex 79.3%
  11. GPT-5.6 Terra 78.3%
  12. Claude Opus 4.6 78%
  13. GPT-5.4 77.7%
  14. Claude Opus 4.7 76.4%
  15. GLM-5.3 75.4%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

WeirdML results by model
#ModelProviderScoreSettingSourceDate
1GPT-6 Astra OpenAI93.6%promaxEpoch AI
2Claude Fable 5.1 Anthropic92.9%maxEpoch AI
3Claude Fable 5 Anthropic91.9%maxEpoch AI
4Claude Opus 5 Anthropic91.8%maxEpoch AI
5GPT-5.6 Sol OpenAI89.4%promaxEpoch AI
6GPT-5.5 OpenAI84.9%xhighEpoch AI
7Gemini 3.8 Flash Google84.8%highEpoch AI
8Claude Opus 4.8 Anthropic82.9%xhighEpoch AI
9Kimi K3 Moonshot AI82.6%maxEpoch AI
10GPT-5.3 Codex OpenAI79.3%Epoch AI
11GPT-5.6 Terra OpenAI78.3%highEpoch AI
12Claude Opus 4.6 Anthropic78%highEpoch AI
13GPT-5.4 OpenAI77.7%xhighEpoch AI
14Claude Opus 4.7 Anthropic76.4%highEpoch AI
15GLM-5.3 Z.ai (Zhipu)75.4%maxEpoch AI
16GPT-5.2 OpenAI72.2%xhighEpoch AI
17Gemini 3.1 Pro Preview Google72.1%Epoch AI
18GLM-5.2 Z.ai (Zhipu)70.1%maxEpoch AI
19Gemini 3 Pro Google69.9%Epoch AI
20Claude Sonnet 5 Anthropic68.8%highEpoch AI
21Grok 4.6 xAI67.3%highEpoch AI
22DeepSeek V4 Pro DeepSeek66.2%maxEpoch AI
23Claude Sonnet 4.6 Anthropic66.1%mediumEpoch AI
24Claude Opus 4.5 Anthropic63.7%16KEpoch AI
25DeepSeek V4 Flash DeepSeek63%maxEpoch AI
26Gemini 3.5 Flash Google62.6%highEpoch AI
27Gemini 3 Flash Preview Google61.6%Epoch AI
28GPT-5.6 Luna OpenAI60.9%highEpoch AI
29GPT-5.1 OpenAI60.8%highEpoch AI
30GPT-5 OpenAI60.7%highEpoch AI
31GPT-5 Pro OpenAI60.4%highEpoch AI
32GPT-5.4 mini OpenAI60.3%highEpoch AI
33Muse Spark 1.2 Meta60.3%xhighEpoch AI
34o3-pro OpenAI58.2%highEpoch AI
35GPT-5.4 Pro OpenAI57.4%noneEpoch AI
36GLM-5.1 Z.ai (Zhipu)57.1%Epoch AI
37Gemini 3.6 Flash Google56.1%highEpoch AI
38Kimi K2.6 Moonshot AI55.9%Epoch AI
39GPT-5-Codex OpenAI54.5%Epoch AI
40Kimi K2.7 Code Moonshot AI54.1%Epoch AI
41Gemini 2.5 Pro Google54%16KEpoch AI
42GPT-5 Mini OpenAI52.7%highEpoch AI
43o4-mini OpenAI52.6%highEpoch AI
44o3 OpenAI52.4%highEpoch AI
45Gemma 4 31B IT Google52.3%Epoch AI
46Grok 4.20 (Non-Reasoning) xAI52.3%Epoch AI
47Gemini 3.1 Flash Lite Google52.2%Epoch AI
48Grok 4.3 xAI49.9%Epoch AI
49GPT-5.4 nano OpenAI49.2%highEpoch AI
50GLM-5 Z.ai (Zhipu)48.2%Epoch AI
51gpt-oss-120b OpenAI48.2%highEpoch AI
52gpt-oss-120b OpenAI48.2%highEpoch AI
53Claude Sonnet 4.5 Anthropic47.7%16KEpoch AI
54o1 OpenAI47.6%Epoch AI
55DeepSeek-V3.2-Speciale DeepSeek46.7%Epoch AI
56Grok 4.5 xAI46.4%Epoch AI
57Claude Sonnet 4 Anthropic46.1%16KEpoch AI
58Claude Opus 4.1 Anthropic45.9%16KEpoch AI
59Grok 4 xAI45.7%Epoch AI
60Kimi K2.5 Moonshot AI45.6%Epoch AI
61Claude Haiku 4.5 Anthropic45.4%Epoch AI
62Claude Opus 4 Anthropic43.7%16KEpoch AI
63Mistral Medium Mistral AI43.7%Epoch AI
64o3-mini OpenAI43.7%highEpoch AI
65Nemotron 3 Ultra NVIDIA43.5%Epoch AI
66Mercury 2 Inception43.2%Epoch AI
67Grok 4 Fast xAI42.9%Epoch AI
68Kimi K2 (Jul 2025) Moonshot AI42.8%Epoch AI
69Grok-3 mini xAI42.6%highEpoch AI
70Grok-3 mini xAI42.6%highEpoch AI
71Gemini 2.5 Flash Google41.9%16kEpoch AI
72DeepSeek-R1 DeepSeek41.6%Epoch AI
73Qwen3-Coder 480B-A35B Instruct Alibaba (Qwen)41.2%Epoch AI
74Qwen3 235B-A22B Alibaba (Qwen)41%Epoch AI
75gpt-oss-20b OpenAI40.9%highEpoch AI
76GLM-4.5 Z.ai (Zhipu)40.6%thinkingEpoch AI
77Claude 3.5 Sonnet Anthropic40%Epoch AI
78Qwen3.5 27B Alibaba (Qwen)39.5%Epoch AI
79DeepSeek-V3.2-Exp DeepSeek39.5%thinkingEpoch AI
80GPT-4.5 OpenAI39.4%Epoch AI
81GPT-4.1 OpenAI39%Epoch AI
82Gemini 3.5 Flash Lite Google39%highEpoch AI
83DeepSeek-V3.1 DeepSeek38.4%thinkingEpoch AI
84GPT-5 Nano OpenAI38.1%highEpoch AI
85Nemotron 3 Super NVIDIA38%Epoch AI
86GPT-4.1 mini OpenAI37.6%Epoch AI
87Grok 3 xAI37.2%Epoch AI
88MiniMax-M2.7 MiniMax37%Epoch AI
89o1-mini OpenAI36.3%mediumEpoch AI
90DeepSeek-V3 DeepSeek36.1%Epoch AI
91Gemini 2.5 Flash-Lite Google35.2%16KEpoch AI
92Gemma 4 26B A4B IT Google35.2%Epoch AI
93Qwen3.6 35B-A3B Alibaba (Qwen)34.5%Epoch AI
94Qwen3 Coder Next Alibaba (Qwen)34.4%Epoch AI
95Inkling Thinking Machines Lab32.3%highEpoch AI
96Claude 3.5 Haiku Anthropic30.7%Epoch AI
97Qwen3-30B-A3B Alibaba (Qwen)29.8%Epoch AI
98Gemini 2.0 Flash (Feb 2025) Google25.8%Epoch AI
99GPT-4o OpenAI25.1%Epoch AI
100Gemini 1.5 Flash (May 2024) Google24.9%Epoch AI
101Llama 4 Maverick Meta24.5%Epoch AI
102Grok-2 (Dec 2024) xAI22.2%Epoch AI
103Gemini 1.5 Pro (May 2024) Google22.2%Epoch AI
104Llama 3.1-405B Meta21.4%Epoch AI
105Claude 3 Opus Anthropic19.2%Epoch AI
106GPT-4.1 nano OpenAI19%Epoch AI
107GPT-4 Turbo OpenAI18%Epoch AI
108Qwen2.5 72B Instruct Alibaba (Qwen)16%Epoch AI
109Llama-3.3-70B-Instruct Meta14.4%Epoch AI
110GPT-4 OpenAI12.4%Epoch AI
111GPT-4o mini OpenAI11.8%Epoch AI
112Qwen2-72B Alibaba (Qwen)11.3%Epoch AI
113Claude 3 Sonnet Anthropic10.2%Epoch AI
114Claude 3 Haiku Anthropic9.8%Epoch AI
115Llama 3.1-70B Meta9%Epoch AI
116Claude 2.1 Anthropic7.1%Epoch AI
117GPT-3.5-turbo OpenAI3.5%Epoch AI
118Mixtral 8x22B Mistral AI3.2%Epoch AI
119Llama 3.1-8B Meta1.7%Epoch AI

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does WeirdML measure?

Unusual machine-learning tasks the model must solve by writing and iterating on working PyTorch code.

Which model has the highest WeirdML score?

As of October 2026, GPT-6 Astra has the highest published WeirdML score on Noometry at 93.6%, out of 119 models with results.

What is the best open-weight model on WeirdML?

Kimi K3 has the highest WeirdML accuracy among open-weight models at 82.6%, ranking 9 of 119 overall.