Model comparison

GPT-4o mini vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 25.5 on the Noometry Index.

Last verified . 21 shared benchmarks.

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GPT-4o mini scores higher in 0 categories and Grok 4.1 Fast in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 8.7.
  • The biggest single-benchmark swing is SimpleBench: 10.7% for GPT-4o mini and 56% for Grok 4.1 Fast.
  • Both cost about the same: $0.15 input and $0.60 output per million tokens.

Side by side

GPT-4o mini and Grok 4.1 Fast specifications
GPT-4o miniGrok 4.1 Fast
ProviderOpenAIxAI
Noometry Index25.541.4
Released2024-07-182025-06-27
WeightsProprietaryProprietary
Context window128K128K
Max output16K30K
Input $ / M tokens$0.15$0.20
Output $ / M tokens$0.60$0.50
Results tracked6032

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 Fast leads

GPT-4o mini: 22.0 (#335), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Coding12901411
Aider Polyglot3.6%—
LMArena WebDev—1242
WeirdML11.8%—
BigCodeBench Instruct46.1%—
LiveBench Coding43.1%—
BigCodeBench Complete57.4%—
ALE-Bench—394.93
HumanEval+83.5%—
MBPP+72.2%—

Agentic & Tool Use Grok 4.1 Fast leads

GPT-4o mini: 27.5 (#101), Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
BALROG17.4%—
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

GPT-4o mini: 8.7 (#347), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
SimpleBench10.7%56%
LMArena Hard Prompts12671407
DTBench54.4%87.7%
ARC-AGI-20%—
Kagi LLM Benchmark28.8%—
NYT Connections (extended)—87.4%
Chess Puzzles0%—
LiveBench Reasoning32.8%—
Mystery Game Puzzles12%—
LiveBench Data Analysis50%—
LMCA10.4%—
Epoch Capabilities Index126.56—
ForecastBench—61
LiveBench41.3%—
PIQA88.7%—

Math Grok 4.1 Fast leads

GPT-4o mini: 10.4 (#314), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Math12671408
FrontierMath (Tiers 1-3)0.7%—
MathArena Final-Answer Competitions—60.9%
OTIS Mock AIME 2024-20256.9%—
ProofBench—4%
Omni-MATH28%—
LiveBench Math36.3%—
MATH Level 552.6%—
GSM8K91.3%—

Knowledge Grok 4.1 Fast leads

GPT-4o mini: 17.7 (#284), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Expert12351399
GPQA Diamond37.7%—
SimpleQA Verified8.3%—
MMLU-Pro60.3%—
Confabulations37.2%—
Vectara Hallucination Rate—17.8%
GPQA (HELM)36.8%—
BoolQ88.7%—
MMLU81.8%—

Multimodal Grok 4.1 Fast leads

GPT-4o mini: 25.9 (#122), Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Vision10661201
Video-MME64.8%—
GeoBench64%—
VPCT34%—

Multilingual Grok 4.1 Fast leads

GPT-4o mini: 42.0 (#199), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Non-English12661391
LMArena Chinese12651441
LMArena French12971415
LMArena German12721404
LMArena Japanese12161349
LMArena Korean11951361
LMArena Russian12751387
LMArena Spanish12761413

Instruction Following Grok 4.1 Fast leads

GPT-4o mini: 61.9 (#239), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Instruction Following12581376
LiveBench Instruction Following56.8%—
IFEval78.2%—

Long Context Grok 4.1 Fast leads

GPT-4o mini: 39.1 (#186), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Longer Query12891390

Writing & Preference Grok 4.1 Fast leads

GPT-4o mini: 39.5 (#248), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGPT-4o miniGrok 4.1 Fast
LMArena Text12861408
LMArena Creative Writing12681394
EQ-Bench Creative Writing8731327
LMArena Multi-Turn12851389
Short-Story Creative Writing67.2%—
WildBench79.1%—
LiveBench Language28.6%—

Frequently asked questions

Is GPT-4o mini better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 25.5 on the Noometry Index.

Which is cheaper, GPT-4o mini or Grok 4.1 Fast?

GPT-4o mini is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Is GPT-4o mini or Grok 4.1 Fast better for coding?

Grok 4.1 Fast scores higher on coding benchmarks: 34.1 versus 22.0 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do GPT-4o mini and Grok 4.1 Fast share?

21 benchmarks have published results for both models. GPT-4o mini has 60 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper