Model comparison

GPT-5.1 vs Grok 4

GPT-5.1 and Grok 4 score almost the same on the Noometry Index (49.0 vs 48.1), so choose on price, context window or the category you care about most.

Last verified . 37 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

Grok 4 xAI

48.1

Rank #56 Confirmed

Summary

  • They share 37 benchmarks with published results for both. GPT-5.1 scores higher in 7 categories and Grok 4 in 3 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok 4 leads 63.1 to 47.6.
  • The biggest single-benchmark swing is GPQA (HELM): 44.2% for GPT-5.1 and 72.7% for Grok 4.

Side by side

GPT-5.1 and Grok 4 specifications
GPT-5.1Grok 4
ProviderOpenAIxAI
Noometry Index49.048.1
Released2025-11-132025-07-09
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$1.25—
Output $ / M tokens$10—
Results tracked6348

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4 leads

GPT-5.1: 46.4 (#66), Grok 4: 50.3 (#46)

Coding benchmarks
BenchmarkGPT-5.1Grok 4
WeirdML60.8%45.7%
LMArena Coding14541408
SWE-bench Verified68%—
SWE-bench Verified (bash only)66%—
Aider Polyglot—79.6%
LMArena WebDev1395—
SciCode43.3%—
GSO13.7%—
LiveBench Coding72.5%—
ALE-Bench1,192—

Agentic & Tool Use Too close to call

GPT-5.1: 32.7 (#60), Grok 4: 32.3 (#68)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1Grok 4
Terminal-Bench47.6%27.2%
DeepResearch Bench42.8%47.3%
LMArena Search11991142
Berkeley Function Calling Leaderboard—63%
GDPval—21.1%
Cybench—43%
BALROG—43.6%
METR Time Horizons—66.6%
Vending-Bench 21,473—

Reasoning GPT-5.1 leads

GPT-5.1: 39.8 (#58), Grok 4: 36.7 (#65)

Reasoning benchmarks
BenchmarkGPT-5.1Grok 4
ARC-AGI-217.6%16%
SimpleBench53.2%60.5%
ARC-AGI-172.8%66.7%
Chess Puzzles32%28%
LMArena Hard Prompts14571409
Epoch Capabilities Index149.64146.44
ForecastBench58.160.9
Kagi LLM Benchmark—73.6%
CritPt4.9%—
EnigmaEval11.2%—
LiveBench Reasoning95.8%—
Mystery Game Puzzles19%—
DTBench90.1%—
LiveBench Data Analysis72.1%—
LMCA43.9%—
LiveBench78.8%—

Math GPT-5.1 leads

GPT-5.1: 52.2 (#51), Grok 4: 48.4 (#64)

Math benchmarks
BenchmarkGPT-5.1Grok 4
OTIS Mock AIME 2024-202588.6%84%
Omni-MATH46.4%60.3%
LMArena Math14471422
FrontierMath (Feb 2025 set)31%19.7%
FrontierMath Tier 4 (v1)12.5%2.1%
LiveBench Math94.5%—

Knowledge Grok 4 leads

GPT-5.1: 50.6 (#71), Grok 4: 53.8 (#55)

Knowledge benchmarks
BenchmarkGPT-5.1Grok 4
GPQA Diamond87.6%87%
MMLU-Pro57.9%85.1%
GPQA (HELM)44.2%72.7%
LMArena Expert14701415
Humanity's Last Exam23.7%—
SimpleQA Verified48%—
Confabulations—12.4%
Vectara Hallucination Rate10.9%—

Multimodal GPT-5.1 leads

GPT-5.1: 44.8 (#19), Grok 4: 33.7 (#94)

Multimodal benchmarks
BenchmarkGPT-5.1Grok 4
LMArena Vision12501210
GeoBench—45%
VPCT58.7%—
LMArena Document1403—

Multilingual GPT-5.1 leads

GPT-5.1: 53.8 (#56), Grok 4: 51.8 (#103)

Multilingual benchmarks
BenchmarkGPT-5.1Grok 4
LMArena Non-English14311403
LMArena Chinese14951427
LMArena French14501418
LMArena German14381429
LMArena Japanese14531394
LMArena Korean14011377
LMArena Russian14351410
LMArena Spanish14331420

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), Grok 4: 79.2 (#5)

Instruction Following benchmarks
BenchmarkGPT-5.1Grok 4
IFEval93.5%94.9%
LMArena Instruction Following14431387
LiveBench Instruction Following93.3%—

Long Context Grok 4 leads

GPT-5.1: 47.6 (#14), Grok 4: 63.1 (#4)

Long Context benchmarks
BenchmarkGPT-5.1Grok 4
LMArena Longer Query14471409
Fiction.LiveBench—94.4%
CL-bench23.7%—
CL-bench Life17.3%—

Writing & Preference GPT-5.1 leads

GPT-5.1: 64.5 (#55), Grok 4: 58.5 (#116)

Writing & Preference benchmarks
BenchmarkGPT-5.1Grok 4
LMArena Text14431411
LMArena Creative Writing14271397
WildBench86.3%79.7%
LMArena Multi-Turn14501416
Short-Story Creative Writing—76.9%
LiveBench Language80.2%—

Frequently asked questions

Is GPT-5.1 better than Grok 4?

GPT-5.1 and Grok 4 score almost the same on the Noometry Index (49.0 vs 48.1), so choose on price, context window or the category you care about most.

Is GPT-5.1 or Grok 4 better for coding?

Grok 4 scores higher on coding benchmarks: 50.3 versus 46.4 in the Noometry coding category.

How many benchmarks do GPT-5.1 and Grok 4 share?

37 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and Grok 4 has 48.

Related comparisons

Go deeper