Model comparison

Grok 4.1 Fast vs o3

o3 is the stronger model overall, scoring 47.5 to 41.4 on the Noometry Index. Grok 4.1 Fast costs 13× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 25 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Grok 4.1 Fast scores higher in 2 categories and o3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 33.1.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 69.6% for Grok 4.1 Fast and 63% for o3.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $2 / $8 for o3.
  • o3 accepts more context: 200K tokens versus 128K.

Side by side

Grok 4.1 Fast and o3 specifications
Grok 4.1 Fasto3
ProviderxAIOpenAI
Noometry Index41.447.5
Released2025-06-272025-04-16
WeightsProprietaryProprietary
Context window128K200K
Max output30K100K
Input $ / M tokens$0.20$2
Output $ / M tokens$0.50$8
Results tracked3263

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Grok 4.1 Fast: 34.1 (#245), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Coding14111408
ALE-Bench394.93933.55
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1242—
GSO—8.8%
WeirdML—52.4%
CadEval—74%

Agentic & Tool Use Grok 4.1 Fast leads

Grok 4.1 Fast: 36.3 (#39), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 Fasto3
Berkeley Function Calling Leaderboard69.6%63%
LMArena Search11711144
GDPval—30.8%
τ²-bench Banking13.1%—
DeepResearch Bench—45.2%
OSWorld—23%
METR Time Horizons—65.4%
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGrok 4.1 Fasto3
SimpleBench56%53.1%
LMArena Hard Prompts14071402
DTBench87.7%84.8%
ForecastBench6162.5
ARC-AGI-2—6.5%
Kagi LLM Benchmark—67.6%
NYT Connections (extended)87.4%—
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
LMCA—39.7%
Epoch Capabilities Index—146.86

Math o3 leads

Grok 4.1 Fast: 31.9 (#221), o3: 50.2 (#58)

Knowledge o3 leads

Grok 4.1 Fast: 33.1 (#207), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Expert13991402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
Vectara Hallucination Rate17.8%—
GPQA (HELM)—75.3%

Multimodal o3 leads

Grok 4.1 Fast: 37.0 (#76), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Vision12011214
GeoBench—74%
VPCT—52%

Multilingual Too close to call

Grok 4.1 Fast: 51.0 (#114), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Non-English13911401
LMArena Chinese14411437
LMArena French14151430
LMArena German14041420
LMArena Japanese13491403
LMArena Korean13611370
LMArena Russian13871406
LMArena Spanish14131395

Instruction Following Too close to call

Grok 4.1 Fast: 72.7 (#133), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Instruction Following13761368
IFEval—86.9%

Long Context o3 leads

Grok 4.1 Fast: 42.4 (#126), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Longer Query13901372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Grok 4.1 Fast: 57.2 (#131), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGrok 4.1 Fasto3
LMArena Text14081410
LMArena Creative Writing13941359
EQ-Bench Creative Writing13271676
LMArena Multi-Turn13891405
Short-Story Creative Writing—83.9%
WildBench—86.1%

Frequently asked questions

Is Grok 4.1 Fast better than o3?

o3 is the stronger model overall, scoring 47.5 to 41.4 on the Noometry Index. Grok 4.1 Fast costs 13× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, Grok 4.1 Fast or o3?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; o3 lists at $2 and $8.

Is Grok 4.1 Fast or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

o3 does, with 200K tokens against 128K.

How many benchmarks do Grok 4.1 Fast and o3 share?

25 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper