Model comparison

GPT-4 vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 29.1 on the Noometry Index.

Last verified . 20 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 20 benchmarks with published results for both. GPT-4 scores higher in 0 categories and Grok 4.1 Fast in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 17.8.
  • The biggest single-benchmark swing is DTBench: 62.7% for GPT-4 and 87.7% for Grok 4.1 Fast.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $30 / $60 for GPT-4.
  • Grok 4.1 Fast accepts more context: 128K tokens versus 8K.

Side by side

GPT-4 and Grok 4.1 Fast specifications
GPT-4Grok 4.1 Fast
ProviderOpenAIxAI
Noometry Index29.141.4
Released2023-03-142025-06-27
WeightsProprietaryProprietary
Context window8K128K
Max output8K30K
Input $ / M tokens$30$0.20
Output $ / M tokens$60$0.50
Results tracked3832

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 Fast leads

GPT-4: 31.6 (#283), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Coding12541411
LMArena WebDev—1242
WeirdML12.4%—
BigCodeBench Instruct46%—
BigCodeBench Complete57.2%—
ALE-Bench—394.93
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGPT-4Grok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
METR Time Horizons36.1%—
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

GPT-4: 17.8 (#289), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Hard Prompts12411407
DTBench62.7%87.7%
ForecastBench57.861
SimpleBench—56%
NYT Connections (extended)—87.4%
Chess Puzzles4%—
Mystery Game Puzzles12%—
LMCA17.1%—
BIG-Bench Hard75.1%—
Epoch Capabilities Index125.89—
HellaSwag95.3%—
WinoGrande87.5%—

Math Grok 4.1 Fast leads

GPT-4: 10.8 (#309), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Math12691408
MathArena Final-Answer Competitions—60.9%
OTIS Mock AIME 2024-20251.1%—
ProofBench—4%
MATH Level 523%—
GSM8K92%—

Knowledge Grok 4.1 Fast leads

GPT-4: 18.4 (#282), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Expert12111399
GPQA Diamond35.7%—
Vectara Hallucination Rate—17.8%
MMLU86.4%—
TriviaQA84.8%—

Multimodal Not comparable

GPT-4: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Vision—1201

Multilingual Grok 4.1 Fast leads

GPT-4: 40.6 (#215), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Non-English12461391
LMArena Chinese12421441
LMArena French12831415
LMArena German12511404
LMArena Japanese12091349
LMArena Korean11841361
LMArena Russian12511387
LMArena Spanish12611413

Instruction Following Grok 4.1 Fast leads

GPT-4: 65.3 (#222), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Instruction Following12411376

Long Context Grok 4.1 Fast leads

GPT-4: 37.7 (#212), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Longer Query12441390

Writing & Preference Grok 4.1 Fast leads

GPT-4: 34.9 (#268), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGPT-4Grok 4.1 Fast
LMArena Text12631408
LMArena Creative Writing12441394
EQ-Bench Creative Writing7521327
LMArena Multi-Turn12571389

Frequently asked questions

Is GPT-4 better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 29.1 on the Noometry Index.

Which is cheaper, GPT-4 or Grok 4.1 Fast?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; GPT-4 lists at $30 and $60.

Is GPT-4 or Grok 4.1 Fast better for coding?

Grok 4.1 Fast scores higher on coding benchmarks: 34.1 versus 31.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.1 Fast does, with 128K tokens against 8K.

How many benchmarks do GPT-4 and Grok 4.1 Fast share?

20 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper