Model comparison

GPT-4o mini vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 25.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-4o mini scores higher in 0 categories and Grok 4.1 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.1 leads 38.9 to 10.4.

Side by side

GPT-4o mini and Grok 4.1 specifications
GPT-4o miniGrok 4.1
ProviderOpenAIxAI
Noometry Index25.541.5
Released2024-07-182025-11-17
WeightsProprietaryProprietary
Context window128K—
Max output16K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked6019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 leads

GPT-4o mini: 22.0 (#335), Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Coding12901445
Aider Polyglot3.6%—
LMArena WebDev—1214
WeirdML11.8%—
BigCodeBench Instruct46.1%—
LiveBench Coding43.1%—
BigCodeBench Complete57.4%—
HumanEval+83.5%—
MBPP+72.2%—

Agentic & Tool Use Grok 4.1 leads

GPT-4o mini: 27.5 (#101), Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkGPT-4o miniGrok 4.1
Cybench—39%
BALROG17.4%—

Reasoning Grok 4.1 leads

GPT-4o mini: 8.7 (#347), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Hard Prompts12671435
ARC-AGI-20%—
SimpleBench10.7%—
Kagi LLM Benchmark28.8%—
Chess Puzzles0%—
LiveBench Reasoning32.8%—
Mystery Game Puzzles12%—
DTBench54.4%—
LiveBench Data Analysis50%—
LMCA10.4%—
Epoch Capabilities Index126.56—
LiveBench41.3%—
PIQA88.7%—

Math Grok 4.1 leads

GPT-4o mini: 10.4 (#314), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Math12671422
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.9%—
Omni-MATH28%—
LiveBench Math36.3%—
MATH Level 552.6%—
GSM8K91.3%—

Knowledge Grok 4.1 leads

GPT-4o mini: 17.7 (#284), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Expert12351417
GPQA Diamond37.7%—
SimpleQA Verified8.3%—
MMLU-Pro60.3%—
Confabulations37.2%—
GPQA (HELM)36.8%—
BoolQ88.7%—
MMLU81.8%—

Multimodal Not comparable

GPT-4o mini: 25.9 (#122), Grok 4.1: —

Multimodal benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Vision1066—
Video-MME64.8%—
GeoBench64%—
VPCT34%—

Multilingual Grok 4.1 leads

GPT-4o mini: 42.0 (#199), Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Non-English12661425
LMArena Chinese12651465
LMArena French12971448
LMArena German12721446
LMArena Japanese12161397
LMArena Korean11951407
LMArena Russian12751434
LMArena Spanish12761438

Instruction Following Grok 4.1 leads

GPT-4o mini: 61.9 (#239), Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Instruction Following12581400
LiveBench Instruction Following56.8%—
IFEval78.2%—

Long Context Grok 4.1 leads

GPT-4o mini: 39.1 (#186), Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Longer Query12891416

Writing & Preference Grok 4.1 leads

GPT-4o mini: 39.5 (#248), Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkGPT-4o miniGrok 4.1
LMArena Text12861437
LMArena Creative Writing12681411
LMArena Multi-Turn12851437
Short-Story Creative Writing67.2%—
EQ-Bench Creative Writing873—
WildBench79.1%—
LiveBench Language28.6%—

Frequently asked questions

Is GPT-4o mini better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 25.5 on the Noometry Index.

Is GPT-4o mini or Grok 4.1 better for coding?

Grok 4.1 scores higher on coding benchmarks: 33.7 versus 22.0 in the Noometry coding category.

How many benchmarks do GPT-4o mini and Grok 4.1 share?

17 benchmarks have published results for both models. GPT-4o mini has 60 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper