Coding benchmark

SciCode leaderboard

As of October 2026, Claude Opus 5.5 has the highest published SciCode score on Noometry at 66.9%, out of 121 models with results.

Last verified

About SciCode

Scientific research coding problems from physics, chemistry, biology and math, broken into sub-steps.

Category
Coding
Introduced
2024
Format
Code generation
Unit
Percent (random guessing ≈ 0%)
Official site
scicode-bench.github.io

Top 15 models

Top models on SciCode
  1. Claude Opus 5.5 66.9%
  2. Claude Fable 5.1 63.1%
  3. Gemini 4 Argon 61.8%
  4. Claude Fable 5 61%
  5. Claude Sonnet 5.5 61%
  6. MiMo-V2.6-Pro 60.9%
  7. Gemini 3.7 Flash 59.8%
  8. Muse Spark 1.3 59.7%
  9. Kimi K3 59.5%
  10. GLM-5.3 59%
  11. Gemini 3.1 Pro Preview 58.9%
  12. Step 5 Preview 58.9%
  13. Muse Spark 1.1 58.8%
  14. Grok 4.7 57.8%
  15. GPT-6 Sol 57.6%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

SciCode results by model
#ModelProviderScoreSettingSourceDate
1Claude Opus 5.5 Anthropic66.9%maxEpoch AI
2Claude Fable 5.1 Anthropic63.1%maxEpoch AI
3Gemini 4 Argon Google61.8%highEpoch AI
4Claude Fable 5 Anthropic61%maxEpoch AI
5Claude Sonnet 5.5 Anthropic61%maxEpoch AI
6MiMo-V2.6-Pro Xiaomi60.9%Epoch AI
7Gemini 3.7 Flash Google59.8%mediumEpoch AI
8Muse Spark 1.3 Meta59.7%xhighEpoch AI
9Kimi K3 Moonshot AI59.5%maxEpoch AI
10GLM-5.3 Z.ai (Zhipu)59%maxEpoch AI
11Gemini 3.1 Pro Preview Google58.9%Epoch AI
12Step 5 Preview StepFun58.9%Epoch AI
13Muse Spark 1.1 Meta58.8%Epoch AI
14Grok 4.7 xAI57.8%highEpoch AI
15GPT-6 Sol OpenAI57.6%maxEpoch AI
16GPT-5.6 Sol OpenAI57.1%maxEpoch AI
17Gemini 3.8 Flash Google56.6%highEpoch AI
18GPT-5.4 OpenAI56.6%xhighEpoch AI
19GPT-6 Astra OpenAI56.5%maxEpoch AI
20Grok 4.6 xAI56.5%highEpoch AI
21Claude Opus 5 Anthropic56.4%maxEpoch AI
22Muse Spark 1.2 Meta56.4%xhighEpoch AI
23GPT-5.5 OpenAI56.1%xhighEpoch AI
24GPT-6.1 Sol OpenAI55.8%highEpoch AI
25GPT-5.6 Terra OpenAI55%maxEpoch AI
26GPT-6 Luna OpenAI54.6%maxEpoch AI
27Claude Opus 4.7 Anthropic54.5%maxEpoch AI
28Claude Sonnet 5 Anthropic54.3%highEpoch AI
29Grok 4.5 xAI54.1%highEpoch AI
30GPT-5.6 Luna OpenAI53.6%maxEpoch AI
31Claude Opus 4.8 Anthropic53.5%maxEpoch AI
32Kimi K2.6 Moonshot AI53.5%Epoch AI
33Qwen3.8 Max Alibaba (Qwen)53.2%Epoch AI
34Gemini 3.5 Flash Google53.1%highEpoch AI
35Gemini 3.6 Flash Google52.7%highEpoch AI
36DeepSeek V4.1 Flash DeepSeek51.9%maxEpoch AI
37GLM-5.3-Flash Z.ai (Zhipu)51.6%Epoch AI
38Muse Spark Meta51.5%Epoch AI
39MiMo-V2.6-Flash Xiaomi51.3%Epoch AI
40DeepSeek V4 Pro DeepSeek51%maxEpoch AI
41GLM-5.2 Z.ai (Zhipu)50.5%maxEpoch AI
42Grok Build 0.1 xAI50.2%Epoch AI
43MiMo-V2.5-Pro Xiaomi50.2%Epoch AI
44DeepSeek V4 Flash DeepSeek49.9%maxEpoch AI
45GPT-5.4 mini OpenAI49.9%xhighEpoch AI
46Kimi K2.5 Moonshot AI49%Epoch AI
47Qwen3.7 Max Alibaba (Qwen)48.8%Epoch AI
48Inkling-Small Thinking Machines Lab48.7%Epoch AI
49GPT-5.5 Instant OpenAI48.6%Epoch AI
50Kimi K2.7 Code Moonshot AI47.5%Epoch AI
51Grok 4.3 xAI47.3%highEpoch AI
52MiniMax-M3 MiniMax47.1%Epoch AI
53Inkling Thinking Machines Lab47%xhighEpoch AI
54MiniMax-M2.7 MiniMax47%Epoch AI
55GPT-5.4 nano OpenAI46.9%xhighEpoch AI
56Claude Sonnet 4.6 Anthropic46.8%maxEpoch AI
57Qwen3.8 27B Alibaba (Qwen)46.6%xhighEpoch AI
58Qwen3.7 Plus Alibaba (Qwen)45.5%Epoch AI
59GLM-4.7 Z.ai (Zhipu)45.1%Epoch AI
60Muse Glimmer Meta44.9%highEpoch AI
61Claude Sonnet 4.5 Anthropic44.7%Epoch AI
62GLM-5.1 Z.ai (Zhipu)43.8%Epoch AI
63Gemma 4 31B IT Google43.4%Epoch AI
64Claude Haiku 4.5 Anthropic43.3%Epoch AI
65GPT-5.1 OpenAI43.3%Epoch AI
66MiMo-V2.5 Xiaomi43.1%Epoch AI
67GPT-5 OpenAI42.9%Epoch AI
68Gemini 2.5 Pro Google42.8%Epoch AI
69Nova 2.0 Pro Preview Amazon42.7%mediumEpoch AI
70Qwen3 235B-A22B Alibaba (Qwen)42.4%Epoch AI
71Ring-2.6-1T Ant Group (inclusionAI)42.4%Epoch AI
72Gemini 3.1 Flash Lite Google41.9%Epoch AI
73Gemini 3.5 Flash Lite Google41.3%Epoch AI
74Qwen3.6 Plus Alibaba (Qwen)40.7%Epoch AI
75DeepSeek-V3.1-Terminus DeepSeek40.6%Epoch AI
76GPT-4.1 mini OpenAI40.4%Epoch AI
77Nemotron 3 Ultra NVIDIA40.3%Epoch AI
78Mistral Medium Mistral AI40.2%Epoch AI
79Claude Sonnet 4 Anthropic40%Epoch AI
80Gemma 4 26B A4B IT Google40%Epoch AI
81Step 3.7 Flash StepFun40%Epoch AI
82o3-mini OpenAI39.8%highEpoch AI
83GPT-5 Mini OpenAI39.2%Epoch AI
84Magistral Medium Mistral AI39.2%Epoch AI
85DeepSeek-V3.2-Exp DeepSeek38.9%thinkingEpoch AI
86Mercury 2 Inception38.7%Epoch AI
87Mercury 2.5 Inception38.5%Epoch AI
88GLM-4.6 Z.ai (Zhipu)38.4%Epoch AI
89Command A Plus Cohere37.8%Epoch AI
90Qwen3.6 27B Alibaba (Qwen)37.3%noneEpoch AI
91Mistral Large Mistral AI36.2%Epoch AI
92Trinity Large Thinking Arcee AI36.1%Epoch AI
93gpt-oss-120b OpenAI36%lowEpoch AI
94Nemotron 3 Super NVIDIA36%Epoch AI
95DeepSeek-V3 DeepSeek35.8%Epoch AI
96Qwen3.6 35B-A3B Alibaba (Qwen)35.8%Epoch AI
97DeepSeek-R1 DeepSeek35.7%Epoch AI
98Qwen3.5 122B-A10B Alibaba (Qwen)35.6%noneEpoch AI
99Qwen3 32B Alibaba (Qwen)35.4%Epoch AI
100Magistral Small Mistral AI35.2%Epoch AI
101gpt-oss-20b OpenAI34.4%highEpoch AI
102Qwen3-30B-A3B Alibaba (Qwen)33.3%Epoch AI
103Llama 4 Maverick Meta33.1%Epoch AI
104Qwen3 Coder Next Alibaba (Qwen)32.3%Epoch AI
105Qwen3 14B Alibaba (Qwen)31.6%Epoch AI
106Qwen3.5 35B-A3B Alibaba (Qwen)29.3%noneEpoch AI
107Devstral Small 2505 Mistral AI28.8%Epoch AI
108Qwen3.5-9B Alibaba (Qwen)27.5%Epoch AI
109Claude 3.5 Haiku Anthropic27.4%Epoch AI
110Mistral Small Mistral AI26.5%Epoch AI
111Llama-3.3-70B-Instruct Meta26%Epoch AI
112GPT-4.1 nano OpenAI25.9%Epoch AI
113MiMo-V2-Flash Xiaomi25.9%Epoch AI
114Granite 4.1 30B IBM25.8%Epoch AI
115Solar Pro 3 Upstage24.7%Epoch AI
116Qwen3 8B Alibaba (Qwen)22.6%Epoch AI
117Gemma 3 27B Google21.2%Epoch AI
118Gemma 3 12B Google17.4%Epoch AI
119Llama 4 Scout Meta17%Epoch AI
120Llama 3.1-8B Meta13.2%Epoch AI
121Phi-4 Mini Microsoft10.8%Epoch AI

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does SciCode measure?

Scientific research coding problems from physics, chemistry, biology and math, broken into sub-steps.

Which model has the highest SciCode score?

As of October 2026, Claude Opus 5.5 has the highest published SciCode score on Noometry at 66.9%, out of 121 models with results.

What is the best open-weight model on SciCode?

MiMo-V2.6-Pro has the highest SciCode accuracy among open-weight models at 60.9%, ranking 6 of 121 overall.