Model comparison

Claude Opus 4.1 vs Grok-2 (Dec 2024)

Claude Opus 4.1 is the stronger model overall, scoring 41.0 to 33.7 on the Noometry Index.

Last verified . 26 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Claude Opus 4.1 scores higher in 8 categories and Grok-2 (Dec 2024) in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Opus 4.1 leads 32.2 to 16.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 68.9% for Claude Opus 4.1 and 11.5% for Grok-2 (Dec 2024).

Side by side

Claude Opus 4.1 and Grok-2 (Dec 2024) specifications
Claude Opus 4.1Grok-2 (Dec 2024)
ProviderAnthropicxAI
Noometry Index41.033.7
Released2025-08-052024-08-13
WeightsProprietaryProprietary
Context window200K—
Max output32K—
Input $ / M tokens$15—
Output $ / M tokens$75—
Results tracked4834

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.1 leads

Claude Opus 4.1: 44.4 (#73), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
WeirdML45.9%22.2%
LMArena Coding14791287
SWE-bench Verified73.3%—
LMArena WebDev1390—
LiveBench Coding—46.4%
ALE-Bench674.77—
AlgoTune1.34—

Agentic & Tool Use Not comparable

Claude Opus 4.1: 35.0 (#41), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
Terminal-Bench38%—
GDPval43.6%—
Cybench42%—
DeepResearch Bench48.3%—
LMArena Search1148—
METR Time Horizons66.8%—

Reasoning Claude Opus 4.1 leads

Claude Opus 4.1: 32.2 (#76), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
SimpleBench60%22.7%
LMArena Hard Prompts14431272
DTBench80%65.2%
Epoch Capabilities Index144.12130.48
Chess Puzzles7%—
EnigmaEval7.2%—
EBR-Bench7.9%—
LiveBench Reasoning—54.8%
Mystery Game Puzzles21%—
LiveBench Data Analysis—54.5%
LMCA37.1%—
ForecastBench62—
LiveBench—54.3%

Math Claude Opus 4.1 leads

Claude Opus 4.1: 22.3 (#277), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
OTIS Mock AIME 2024-202568.9%11.5%
LMArena Math14311283
FrontierMath (Feb 2025 set)7.2%0.7%
FrontierMath (Tiers 1-3)12.6%—
FrontierMath Tier 42.4%—
LiveBench Math—54.9%
MATH Level 5—63.5%
FrontierMath Tier 4 (v1)4.2%—

Knowledge Claude Opus 4.1 leads

Claude Opus 4.1: 42.0 (#101), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
GPQA Diamond77.3%53.8%
Confabulations17.1%20.1%
LMArena Expert14391254
Humanity's Last Exam11.5%—
Vectara Hallucination Rate11.8%—

Multimodal Not comparable

Claude Opus 4.1: 26.8 (#119), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
VPCT35%—

Multilingual Claude Opus 4.1 leads

Claude Opus 4.1: 52.0 (#95), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
LMArena Non-English14051282
LMArena Chinese14271289
LMArena French14311318
LMArena German14131287
LMArena Japanese13781244
LMArena Korean13801237
LMArena Russian14221286
LMArena Spanish14481281

Instruction Following Claude Opus 4.1 leads

Claude Opus 4.1: 75.6 (#58), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
LMArena Instruction Following14351270
LiveBench Instruction Following—69.6%

Long Context Claude Opus 4.1 leads

Claude Opus 4.1: 44.5 (#63), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
LMArena Longer Query14551276

Writing & Preference Claude Opus 4.1 leads

Claude Opus 4.1: 62.4 (#74), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1Grok-2 (Dec 2024)
LMArena Text14191305
LMArena Creative Writing14121284
Short-Story Creative Writing84.7%63.6%
LMArena Multi-Turn14441290
LiveBench Language—45.6%

Frequently asked questions

Is Claude Opus 4.1 better than Grok-2 (Dec 2024)?

Claude Opus 4.1 is the stronger model overall, scoring 41.0 to 33.7 on the Noometry Index.

Is Claude Opus 4.1 or Grok-2 (Dec 2024) better for coding?

Claude Opus 4.1 scores higher on coding benchmarks: 44.4 versus 33.3 in the Noometry coding category.

How many benchmarks do Claude Opus 4.1 and Grok-2 (Dec 2024) share?

26 benchmarks have published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper