Model comparison

Grok 3 vs o1-mini

Grok 3 is the stronger model overall, scoring 39.9 to 34.0 on the Noometry Index.

Last verified . 29 shared benchmarks.

Grok 3 xAI

39.9

Rank #157 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Grok 3 scores higher in 8 categories and o1-mini in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 3 leads 46.2 to 34.9.
  • The biggest single-benchmark swing is Aider Polyglot: 53.3% for Grok 3 and 32.9% for o1-mini.

Side by side

Grok 3 and o1-mini specifications
Grok 3o1-mini
ProviderxAIOpenAI
Noometry Index39.934.0
Released2025-04-092024-09-12
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4039

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 3 leads

Grok 3: 41.9 (#115), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkGrok 3o1-mini
Aider Polyglot53.3%32.9%
WeirdML37.2%36.3%
LMArena Coding14321362
LiveBench Coding—48%
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use Grok 3 leads

Grok 3: 30.5 (#76), o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkGrok 3o1-mini
Cybench—10%
BALROG29.5%—

Reasoning Grok 3 leads

Grok 3: 13.7 (#333), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkGrok 3o1-mini
ARC-AGI-20%0.8%
SimpleBench36.1%18.1%
ARC-AGI-15.5%14%
LMArena Hard Prompts14341333
Epoch Capabilities Index138.33135.82
Kagi LLM Benchmark61.3%—
LiveBench Reasoning—72.3%
LiveBench Data Analysis—57.9%
LiveBench—57.8%

Math Grok 3 leads

Grok 3: 38.0 (#145), o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkGrok 3o1-mini
OTIS Mock AIME 2024-202555.6%46.9%
LMArena Math13911358
MATH Level 588.7%89.2%
FrontierMath (Feb 2025 set)3.8%1.7%
Omni-MATH46.4%—
LiveBench Math—62%
FrontierMath Tier 4 (v1)0%—

Knowledge Grok 3 leads

Grok 3: 46.2 (#82), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkGrok 3o1-mini
GPQA Diamond75.8%62.4%
Confabulations14.2%18.6%
LMArena Expert14211316
MMLU-Pro78.8%—
Vectara Hallucination Rate5.8%—
GPQA (HELM)65%—

Multilingual Grok 3 leads

Grok 3: 52.3 (#87), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkGrok 3o1-mini
LMArena Non-English14101289
LMArena Chinese14481314
LMArena French14601293
LMArena German14311278
LMArena Japanese13871245
LMArena Korean13731223
LMArena Russian14161283
LMArena Spanish14171303

Instruction Following Grok 3 leads

Grok 3: 75.0 (#73), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkGrok 3o1-mini
LMArena Instruction Following14091304
LiveBench Instruction Following—65.4%
IFEval88.4%—

Long Context o1-mini leads

Grok 3: 38.7 (#192), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkGrok 3o1-mini
LMArena Longer Query14391320
Fiction.LiveBench58.3%—

Writing & Preference Grok 3 leads

Grok 3: 55.8 (#141), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkGrok 3o1-mini
LMArena Text14261317
LMArena Creative Writing14141244
Short-Story Creative Writing76.4%64.9%
LMArena Multi-Turn14251314
EQ-Bench Creative Writing1186—
WildBench84.9%—
LiveBench Language—40.9%

Frequently asked questions

Is Grok 3 better than o1-mini?

Grok 3 is the stronger model overall, scoring 39.9 to 34.0 on the Noometry Index.

Is Grok 3 or o1-mini better for coding?

Grok 3 scores higher on coding benchmarks: 41.9 versus 35.5 in the Noometry coding category.

How many benchmarks do Grok 3 and o1-mini share?

29 benchmarks have published results for both models. Grok 3 has 40 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper