Model comparison

GPT-5.1 vs Grok-2 (Dec 2024)

GPT-5.1 is the stronger model overall, scoring 49.0 to 33.7 on the Noometry Index.

Last verified . 31 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GPT-5.1 scores higher in 8 categories and Grok-2 (Dec 2024) in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.1 leads 52.2 to 20.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.6% for GPT-5.1 and 11.5% for Grok-2 (Dec 2024).

Side by side

GPT-5.1 and Grok-2 (Dec 2024) specifications
GPT-5.1Grok-2 (Dec 2024)
ProviderOpenAIxAI
Noometry Index49.033.7
Released2025-11-132024-08-13
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$1.25—
Output $ / M tokens$10—
Results tracked6334

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1 leads

GPT-5.1: 46.4 (#66), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
WeirdML60.8%22.2%
LiveBench Coding72.5%46.4%
LMArena Coding14541287
SWE-bench Verified68%—
SWE-bench Verified (bash only)66%—
LMArena WebDev1395—
SciCode43.3%—
GSO13.7%—
ALE-Bench1,192—

Agentic & Tool Use Not comparable

GPT-5.1: 32.7 (#60), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
Terminal-Bench47.6%—
DeepResearch Bench42.8%—
LMArena Search1199—
Vending-Bench 21,473—

Reasoning GPT-5.1 leads

GPT-5.1: 39.8 (#58), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
SimpleBench53.2%22.7%
LiveBench Reasoning95.8%54.8%
LMArena Hard Prompts14571272
DTBench90.1%65.2%
LiveBench Data Analysis72.1%54.5%
Epoch Capabilities Index149.64130.48
LiveBench78.8%54.3%
ARC-AGI-217.6%—
ARC-AGI-172.8%—
CritPt4.9%—
Chess Puzzles32%—
EnigmaEval11.2%—
Mystery Game Puzzles19%—
LMCA43.9%—
ForecastBench58.1—

Math GPT-5.1 leads

GPT-5.1: 52.2 (#51), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
OTIS Mock AIME 2024-202588.6%11.5%
LiveBench Math94.5%54.9%
LMArena Math14471283
FrontierMath (Feb 2025 set)31%0.7%
Omni-MATH46.4%—
MATH Level 5—63.5%
FrontierMath Tier 4 (v1)12.5%—

Knowledge GPT-5.1 leads

GPT-5.1: 50.6 (#71), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
GPQA Diamond87.6%53.8%
LMArena Expert14701254
Humanity's Last Exam23.7%—
SimpleQA Verified48%—
MMLU-Pro57.9%—
Confabulations—20.1%
Vectara Hallucination Rate10.9%—
GPQA (HELM)44.2%—

Multimodal Not comparable

GPT-5.1: 44.8 (#19), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
LMArena Vision1250—
VPCT58.7%—
LMArena Document1403—

Multilingual GPT-5.1 leads

GPT-5.1: 53.8 (#56), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
LMArena Non-English14311282
LMArena Chinese14951289
LMArena French14501318
LMArena German14381287
LMArena Japanese14531244
LMArena Korean14011237
LMArena Russian14351286
LMArena Spanish14331281

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
LiveBench Instruction Following93.3%69.6%
LMArena Instruction Following14431270
IFEval93.5%—

Long Context GPT-5.1 leads

GPT-5.1: 47.6 (#14), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
LMArena Longer Query14471276
CL-bench23.7%—
CL-bench Life17.3%—

Writing & Preference GPT-5.1 leads

GPT-5.1: 64.5 (#55), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkGPT-5.1Grok-2 (Dec 2024)
LMArena Text14431305
LMArena Creative Writing14271284
LMArena Multi-Turn14501290
LiveBench Language80.2%45.6%
Short-Story Creative Writing—63.6%
WildBench86.3%—

Frequently asked questions

Is GPT-5.1 better than Grok-2 (Dec 2024)?

GPT-5.1 is the stronger model overall, scoring 49.0 to 33.7 on the Noometry Index.

Is GPT-5.1 or Grok-2 (Dec 2024) better for coding?

GPT-5.1 scores higher on coding benchmarks: 46.4 versus 33.3 in the Noometry coding category.

How many benchmarks do GPT-5.1 and Grok-2 (Dec 2024) share?

31 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper