Model comparison

GPT-4.1 mini vs Grok-3 mini

Grok-3 mini is the stronger model overall, scoring 41.2 to 33.6 on the Noometry Index.

Last verified . 33 shared benchmarks.

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Summary

  • They share 33 benchmarks with published results for both. GPT-4.1 mini scores higher in 0 categories and Grok-3 mini in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok-3 mini leads 42.1 to 24.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 44.7% for GPT-4.1 mini and 77.8% for Grok-3 mini.

Side by side

GPT-4.1 mini and Grok-3 mini specifications
GPT-4.1 miniGrok-3 mini
ProviderOpenAIxAI
Noometry Index33.641.2
Released2025-04-142025-04-09
WeightsProprietaryProprietary
Context window1.05M—
Max output33K—
Input $ / M tokens$0.40—
Output $ / M tokens$1.60—
Results tracked4735

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok-3 mini leads

GPT-4.1 mini: 30.6 (#293), Grok-3 mini: 40.8 (#131)

Coding benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
Aider Polyglot32.4%49.3%
WeirdML37.6%42.6%
LMArena Coding13671379
SWE-bench Verified (bash only)23.9%—
SciCode40.4%—
BigCodeBench Instruct48.9%—
CadEval16%—

Agentic & Tool Use Not comparable

GPT-4.1 mini: 33.3 (#55), Grok-3 mini: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
Berkeley Function Calling Leaderboard50.5%—

Reasoning Grok-3 mini leads

GPT-4.1 mini: 10.8 (#340), Grok-3 mini: 13.6 (#334)

Reasoning benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
ARC-AGI-20%0.4%
Kagi LLM Benchmark48.6%61.3%
ARC-AGI-13.5%16.5%
LMArena Hard Prompts13491375
Epoch Capabilities Index135.01140.35
CritPt0%—
Chess Puzzles7%—
Mystery Game Puzzles7%—
DTBench68.8%—
LMCA21.1%—

Math Grok-3 mini leads

GPT-4.1 mini: 24.1 (#270), Grok-3 mini: 42.1 (#85)

Math benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
OTIS Mock AIME 2024-202544.7%77.8%
Omni-MATH49.1%31.8%
LMArena Math13431386
MATH Level 587.3%90.9%
FrontierMath (Feb 2025 set)4.5%5.9%
FrontierMath (Tiers 1-3)6.7%—

Knowledge Grok-3 mini leads

GPT-4.1 mini: 34.7 (#194), Grok-3 mini: 46.4 (#81)

Knowledge benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
GPQA Diamond65.8%76.3%
MMLU-Pro78.3%79.9%
GPQA (HELM)61.4%67.5%
LMArena Expert13381395
SimpleQA Verified12.7%—
Confabulations—10.8%

Multimodal Not comparable

GPT-4.1 mini: 35.8 (#82), Grok-3 mini: —

Multimodal benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
LMArena Vision1181—

Multilingual Grok-3 mini leads

GPT-4.1 mini: 45.7 (#166), Grok-3 mini: 48.1 (#145)

Multilingual benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
LMArena Non-English13181352
LMArena Chinese13291387
LMArena French13581357
LMArena German13511349
LMArena Japanese12901342
LMArena Korean12981335
LMArena Russian13241353
LMArena Spanish13191381

Instruction Following Grok-3 mini leads

GPT-4.1 mini: 73.7 (#118), Grok-3 mini: 78.5 (#9)

Instruction Following benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
IFEval90.4%95.1%
LMArena Instruction Following13331357

Long Context Grok-3 mini leads

GPT-4.1 mini: 31.8 (#275), Grok-3 mini: 41.0 (#147)

Long Context benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
Fiction.LiveBench44.4%66.7%
LMArena Longer Query13441372

Writing & Preference Grok-3 mini leads

GPT-4.1 mini: 48.6 (#199), Grok-3 mini: 52.5 (#169)

Writing & Preference benchmarks
BenchmarkGPT-4.1 miniGrok-3 mini
LMArena Text13401370
LMArena Creative Writing13001342
WildBench83.8%65.1%
LMArena Multi-Turn13541355
Short-Story Creative Writing—73.5%
EQ-Bench Creative Writing1147—

Frequently asked questions

Is GPT-4.1 mini better than Grok-3 mini?

Grok-3 mini is the stronger model overall, scoring 41.2 to 33.6 on the Noometry Index.

Is GPT-4.1 mini or Grok-3 mini better for coding?

Grok-3 mini scores higher on coding benchmarks: 40.8 versus 30.6 in the Noometry coding category.

How many benchmarks do GPT-4.1 mini and Grok-3 mini share?

33 benchmarks have published results for both models. GPT-4.1 mini has 47 scored results on Noometry and Grok-3 mini has 35.

Related comparisons

Go deeper