Model comparison

Grok 4.1 vs o3

o3 is the stronger model overall, scoring 47.5 to 41.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 4.1 scores higher in 2 categories and o3 in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 39.5.

Side by side

Grok 4.1 and o3 specifications
Grok 4.1o3
ProviderxAIOpenAI
Noometry Index41.547.5
Released2025-11-172025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Grok 4.1: 33.7 (#253), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGrok 4.1o3
LMArena Coding14451408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1214—
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Too close to call

Grok 4.1: 34.1 (#49), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1o3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
Cybench39%—
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Grok 4.1: 29.5 (#91), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGrok 4.1o3
LMArena Hard Prompts14351402
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Grok 4.1: 38.9 (#120), o3: 50.2 (#58)

Math benchmarks
BenchmarkGrok 4.1o3
LMArena Math14221426
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Grok 4.1: 39.5 (#133), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGrok 4.1o3
LMArena Expert14171402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal Not comparable

Grok 4.1: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGrok 4.1o3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGrok 4.1o3
LMArena Non-English14251401
LMArena Chinese14651437
LMArena French14481430
LMArena German14461420
LMArena Japanese13971403
LMArena Korean14071370
LMArena Russian14341406
LMArena Spanish14381395

Instruction Following Grok 4.1 leads

Grok 4.1: 73.8 (#111), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGrok 4.1o3
LMArena Instruction Following14001368
IFEval—86.9%

Long Context o3 leads

Grok 4.1: 43.2 (#100), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGrok 4.1o3
LMArena Longer Query14161372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Grok 4.1: 62.4 (#75), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGrok 4.1o3
LMArena Text14371410
LMArena Creative Writing14111359
LMArena Multi-Turn14371405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Grok 4.1 better than o3?

o3 is the stronger model overall, scoring 47.5 to 41.5 on the Noometry Index.

Is Grok 4.1 or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and o3 share?

17 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper