Model comparison

Grok 3 vs o1

Grok 3 and o1 score almost the same on the Noometry Index (39.9 vs 40.9), so choose on price, context window or the category you care about most.

Last verified . 29 shared benchmarks.

Grok 3 xAI

39.9

Rank #157 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Grok 3 scores higher in 6 categories and o1 in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where o1 leads 27.9 to 13.7.
  • The biggest single-benchmark swing is ARC-AGI-1: 5.5% for Grok 3 and 30.7% for o1.

Side by side

Grok 3 and o1 specifications
Grok 3o1
ProviderxAIOpenAI
Noometry Index39.940.9
Released2025-04-092024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked4052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Grok 3: 41.9 (#115), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGrok 3o1
Aider Polyglot53.3%61.7%
WeirdML37.2%47.6%
LMArena Coding14321367
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Grok 3 leads

Grok 3: 30.5 (#76), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGrok 3o1
Cybench—10%
BALROG29.5%—
METR Time Horizons—51.1%

Reasoning o1 leads

Grok 3: 13.7 (#333), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGrok 3o1
SimpleBench36.1%41.7%
ARC-AGI-15.5%30.7%
LMArena Hard Prompts14341371
Epoch Capabilities Index138.33141.91
ARC-AGI-20%—
Kagi LLM Benchmark61.3%—
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
LiveBench—75.7%

Math Grok 3 leads

Grok 3: 38.0 (#145), o1: 36.1 (#175)

Knowledge Grok 3 leads

Grok 3: 46.2 (#82), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGrok 3o1
GPQA Diamond75.8%76.8%
Confabulations14.2%11.7%
LMArena Expert14211361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
MMLU-Pro78.8%—
Vectara Hallucination Rate5.8%—
GPQA (HELM)65%—

Multimodal Not comparable

Grok 3: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGrok 3o1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Grok 3 leads

Grok 3: 52.3 (#87), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGrok 3o1
LMArena Non-English14101358
LMArena Chinese14481394
LMArena French14601344
LMArena German14311337
LMArena Japanese13871346
LMArena Korean13731396
LMArena Russian14161356
LMArena Spanish14171345

Instruction Following Too close to call

Grok 3: 75.0 (#73), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGrok 3o1
LMArena Instruction Following14091367
LiveBench Instruction Following—81.5%
IFEval88.4%—

Long Context o1 leads

Grok 3: 38.7 (#192), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGrok 3o1
Fiction.LiveBench58.3%83.3%
LMArena Longer Query14391378

Writing & Preference Too close to call

Grok 3: 55.8 (#141), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGrok 3o1
LMArena Text14261366
LMArena Creative Writing14141348
Short-Story Creative Writing76.4%70.2%
LMArena Multi-Turn14251369
EQ-Bench Creative Writing1186—
WildBench84.9%—
LiveBench Language—65.4%

Frequently asked questions

Is Grok 3 better than o1?

Grok 3 and o1 score almost the same on the Noometry Index (39.9 vs 40.9), so choose on price, context window or the category you care about most.

Is Grok 3 or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 41.9 in the Noometry coding category.

How many benchmarks do Grok 3 and o1 share?

29 benchmarks have published results for both models. Grok 3 has 40 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper