Model comparison

GPT-4 Turbo vs Grok-2 (Dec 2024)

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 30.5 on the Noometry Index.

Last verified . 25 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GPT-4 Turbo scores higher in 1 category and Grok-2 (Dec 2024) in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok-2 (Dec 2024) leads 20.8 to 9.0.
  • The biggest single-benchmark swing is MATH Level 5: 46.7% for GPT-4 Turbo and 63.5% for Grok-2 (Dec 2024).

Side by side

GPT-4 Turbo and Grok-2 (Dec 2024) specifications
GPT-4 TurboGrok-2 (Dec 2024)
ProviderOpenAIxAI
Noometry Index30.533.7
Released2023-11-062024-08-13
WeightsProprietaryProprietary
Context window128K—
Max output4K—
Input $ / M tokens$10—
Output $ / M tokens$30—
Results tracked3634

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4 Turbo: 33.8 (#249), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
WeirdML18%22.2%
LMArena Coding12681287
BigCodeBench Instruct48.2%—
LiveBench Coding—46.4%
BigCodeBench Complete58.2%—
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
METR Time Horizons36.7%—

Reasoning Grok-2 (Dec 2024) leads

GPT-4 Turbo: 15.3 (#317), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
SimpleBench25.1%22.7%
LMArena Hard Prompts12511272
DTBench61.6%65.2%
Epoch Capabilities Index127.25130.48
Chess Puzzles6%—
LiveBench Reasoning—54.8%
LiveBench Data Analysis—54.5%
LMCA9.8%—
ForecastBench59.4—
LiveBench—54.3%

Math Grok-2 (Dec 2024) leads

GPT-4 Turbo: 9.0 (#322), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
OTIS Mock AIME 2024-20256.7%11.5%
LMArena Math12721283
MATH Level 546.7%63.5%
FrontierMath (Tiers 1-3)0.7%—
LiveBench Math—54.9%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Grok-2 (Dec 2024) leads

GPT-4 Turbo: 24.3 (#268), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
GPQA Diamond46.6%53.8%
Confabulations28.4%20.1%
LMArena Expert12231254
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
LMArena Vision1090—

Multilingual Grok-2 (Dec 2024) leads

GPT-4 Turbo: 40.5 (#216), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
LMArena Non-English12451282
LMArena Chinese12421289
LMArena French12761318
LMArena German12591287
LMArena Japanese11941244
LMArena Korean11871237
LMArena Russian12591286
LMArena Spanish12601281

Instruction Following Grok-2 (Dec 2024) leads

GPT-4 Turbo: 65.8 (#216), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
LMArena Instruction Following12491270
LiveBench Instruction Following—69.6%

Long Context Too close to call

GPT-4 Turbo: 38.0 (#206), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
LMArena Longer Query12541276

Writing & Preference Too close to call

GPT-4 Turbo: 47.7 (#206), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkGPT-4 TurboGrok-2 (Dec 2024)
LMArena Text12721305
LMArena Creative Writing12691284
LMArena Multi-Turn12671290
Short-Story Creative Writing—63.6%
LiveBench Language—45.6%

Frequently asked questions

Is GPT-4 Turbo better than Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 30.5 on the Noometry Index.

Is GPT-4 Turbo or Grok-2 (Dec 2024) better for coding?

They score almost the same on coding (33.8 vs 33.3); test both on your own repository before choosing.

How many benchmarks do GPT-4 Turbo and Grok-2 (Dec 2024) share?

25 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper