Model comparison

Claude 3 Opus vs Grok-2 (Dec 2024)

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 29.5 on the Noometry Index.

Last verified . 32 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Claude 3 Opus scores higher in 0 categories and Grok-2 (Dec 2024) in 8 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok-2 (Dec 2024) leads 20.8 to 14.8.
  • The biggest single-benchmark swing is MATH Level 5: 37.5% for Claude 3 Opus and 63.5% for Grok-2 (Dec 2024).

Side by side

Claude 3 Opus and Grok-2 (Dec 2024) specifications
Claude 3 OpusGrok-2 (Dec 2024)
ProviderAnthropicxAI
Noometry Index29.533.7
Released2024-02-292024-08-13
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4634

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3 Opus: 32.9 (#267), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
WeirdML19.2%22.2%
LiveBench Coding38.6%46.4%
LMArena Coding12641287
BigCodeBench Instruct45.5%—
BigCodeBench Complete57.4%—
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

Claude 3 Opus: 24.6 (#116), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
Cybench10%—
METR Time Horizons29.5%—

Reasoning Grok-2 (Dec 2024) leads

Claude 3 Opus: 14.6 (#324), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
SimpleBench23.5%22.7%
LiveBench Reasoning40.6%54.8%
LMArena Hard Prompts12451272
DTBench61.6%65.2%
LiveBench Data Analysis57.9%54.5%
Epoch Capabilities Index126.91130.48
LiveBench49.2%54.3%
Chess Puzzles5%—
EnigmaEval0.8%—
LMCA17%—
ForecastBench58.4—
WinoGrande88.5%—

Math Grok-2 (Dec 2024) leads

Claude 3 Opus: 14.8 (#299), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
OTIS Mock AIME 2024-20254.7%11.5%
LiveBench Math43.6%54.9%
LMArena Math12731283
MATH Level 537.5%63.5%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Grok-2 (Dec 2024) leads

Claude 3 Opus: 24.5 (#267), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
GPQA Diamond47.2%53.8%
Confabulations22.7%20.1%
LMArena Expert12231254
SimpleQA Verified12.6%—
MMLU84.6%—

Multimodal Not comparable

Claude 3 Opus: 27.1 (#116), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
LMArena Vision1023—

Multilingual Grok-2 (Dec 2024) leads

Claude 3 Opus: 41.4 (#207), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
LMArena Non-English12581282
LMArena Chinese12481289
LMArena French12751318
LMArena German12581287
LMArena Japanese12041244
LMArena Korean11871237
LMArena Russian12801286
LMArena Spanish12461281

Instruction Following Grok-2 (Dec 2024) leads

Claude 3 Opus: 64.1 (#228), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
LiveBench Instruction Following63.9%69.6%
LMArena Instruction Following12481270

Long Context Too close to call

Claude 3 Opus: 38.2 (#202), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
LMArena Longer Query12591276

Writing & Preference Grok-2 (Dec 2024) leads

Claude 3 Opus: 47.2 (#213), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGrok-2 (Dec 2024)
LMArena Text12621305
LMArena Creative Writing12351284
LMArena Multi-Turn12751290
LiveBench Language50.4%45.6%
Short-Story Creative Writing—63.6%

Frequently asked questions

Is Claude 3 Opus better than Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 29.5 on the Noometry Index.

Is Claude 3 Opus or Grok-2 (Dec 2024) better for coding?

They score almost the same on coding (32.9 vs 33.3); test both on your own repository before choosing.

How many benchmarks do Claude 3 Opus and Grok-2 (Dec 2024) share?

32 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper