Model comparison

Grok-2 (Dec 2024) vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 33.7 on the Noometry Index.

Last verified . 19 shared benchmarks.

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Grok-2 (Dec 2024) scores higher in 0 categories and Grok 4.1 Fast in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 16.9.
  • The biggest single-benchmark swing is SimpleBench: 22.7% for Grok-2 (Dec 2024) and 56% for Grok 4.1 Fast.

Side by side

Grok-2 (Dec 2024) and Grok 4.1 Fast specifications
Grok-2 (Dec 2024)Grok 4.1 Fast
ProviderxAIxAI
Noometry Index33.741.4
Released2024-08-132025-06-27
WeightsProprietaryProprietary
Context window—128K
Max output—30K
Input $ / M tokens—$0.20
Output $ / M tokens—$0.50
Results tracked3432

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok-2 (Dec 2024): 33.3 (#258), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Coding12871411
LMArena WebDev—1242
WeirdML22.2%—
LiveBench Coding46.4%—
ALE-Bench—394.93

Agentic & Tool Use Not comparable

Grok-2 (Dec 2024): —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Grok-2 (Dec 2024): 16.9 (#299), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
SimpleBench22.7%56%
LMArena Hard Prompts12721407
DTBench65.2%87.7%
NYT Connections (extended)—87.4%
LiveBench Reasoning54.8%—
LiveBench Data Analysis54.5%—
Epoch Capabilities Index130.48—
ForecastBench—61
LiveBench54.3%—

Math Grok 4.1 Fast leads

Grok-2 (Dec 2024): 20.8 (#284), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Math12831408
MathArena Final-Answer Competitions—60.9%
OTIS Mock AIME 2024-202511.5%—
ProofBench—4%
LiveBench Math54.9%—
MATH Level 563.5%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Grok 4.1 Fast leads

Grok-2 (Dec 2024): 29.8 (#233), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Expert12541399
GPQA Diamond53.8%—
Confabulations20.1%—
Vectara Hallucination Rate—17.8%

Multimodal Not comparable

Grok-2 (Dec 2024): —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Vision—1201

Multilingual Grok 4.1 Fast leads

Grok-2 (Dec 2024): 43.1 (#188), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Non-English12821391
LMArena Chinese12891441
LMArena French13181415
LMArena German12871404
LMArena Japanese12441349
LMArena Korean12371361
LMArena Russian12861387
LMArena Spanish12811413

Instruction Following Grok 4.1 Fast leads

Grok-2 (Dec 2024): 66.9 (#202), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Instruction Following12701376
LiveBench Instruction Following69.6%—

Long Context Grok 4.1 Fast leads

Grok-2 (Dec 2024): 38.8 (#190), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Longer Query12761390

Writing & Preference Grok 4.1 Fast leads

Grok-2 (Dec 2024): 48.6 (#198), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGrok-2 (Dec 2024)Grok 4.1 Fast
LMArena Text13051408
LMArena Creative Writing12841394
LMArena Multi-Turn12901389
Short-Story Creative Writing63.6%—
EQ-Bench Creative Writing—1327
LiveBench Language45.6%—

Frequently asked questions

Is Grok-2 (Dec 2024) better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 33.7 on the Noometry Index.

Is Grok-2 (Dec 2024) or Grok 4.1 Fast better for coding?

They score almost the same on coding (33.3 vs 34.1); test both on your own repository before choosing.

How many benchmarks do Grok-2 (Dec 2024) and Grok 4.1 Fast share?

19 benchmarks have published results for both models. Grok-2 (Dec 2024) has 34 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper