Model comparison

Claude Opus 4 vs Grok 4

Grok 4 is the stronger model overall, scoring 48.1 to 43.1 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Grok 4 xAI

48.1

Rank #56 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Opus 4 scores higher in 2 categories and Grok 4 in 8 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok 4 leads 63.1 to 39.6.
  • The biggest single-benchmark swing is Fiction.LiveBench: 61.1% for Claude Opus 4 and 94.4% for Grok 4.

Side by side

Claude Opus 4 and Grok 4 specifications
Claude Opus 4Grok 4
ProviderAnthropicxAI
Noometry Index43.148.1
Released2025-05-222025-07-09
WeightsProprietaryProprietary
Context window200K—
Max output32K—
Input $ / M tokens$15—
Output $ / M tokens$75—
Results tracked5648

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4 leads

Claude Opus 4: 47.2 (#62), Grok 4: 50.3 (#46)

Coding benchmarks
BenchmarkClaude Opus 4Grok 4
Aider Polyglot72%79.6%
WeirdML43.7%45.7%
LMArena Coding14421408
SWE-bench Verified70.7%—
SWE-bench Verified (bash only)67.6%—
GSO6.9%—
AlgoTune1.33—

Agentic & Tool Use Claude Opus 4 leads

Claude Opus 4: 34.8 (#42), Grok 4: 32.3 (#68)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Grok 4
Cybench38%43%
DeepResearch Bench46.8%47.3%
LMArena Search11271142
METR Time Horizons63.9%66.6%
Terminal-Bench—27.2%
Berkeley Function Calling Leaderboard—63%
GDPval—21.1%
BALROG—43.6%

Reasoning Grok 4 leads

Claude Opus 4: 27.3 (#121), Grok 4: 36.7 (#65)

Reasoning benchmarks
BenchmarkClaude Opus 4Grok 4
ARC-AGI-28.6%16%
SimpleBench58.8%60.5%
Kagi LLM Benchmark74.3%73.6%
ARC-AGI-135.7%66.7%
LMArena Hard Prompts13991409
Epoch Capabilities Index142.67146.44
ForecastBench61.160.9
CritPt0.3%—
Chess Puzzles—28%
EnigmaEval5.6%—
DTBench81.6%—
LMCA37.4%—

Math Grok 4 leads

Claude Opus 4: 42.0 (#86), Grok 4: 48.4 (#64)

Math benchmarks
BenchmarkClaude Opus 4Grok 4
OTIS Mock AIME 2024-202564.4%84%
Omni-MATH61.6%60.3%
LMArena Math13901422
FrontierMath (Feb 2025 set)4.5%19.7%
FrontierMath Tier 4 (v1)4.2%2.1%
MATH Level 585%—

Knowledge Grok 4 leads

Claude Opus 4: 44.0 (#88), Grok 4: 53.8 (#55)

Knowledge benchmarks
BenchmarkClaude Opus 4Grok 4
GPQA Diamond76.3%87%
MMLU-Pro87.5%85.1%
Confabulations15.9%12.4%
GPQA (HELM)70.8%72.7%
LMArena Expert13861415
Humanity's Last Exam10.7%—
Vectara Hallucination Rate12%—

Multimodal Grok 4 leads

Claude Opus 4: 31.5 (#106), Grok 4: 33.7 (#94)

Multimodal benchmarks
BenchmarkClaude Opus 4Grok 4
LMArena Vision11921210
GeoBench49%45%
VPCT38%—

Multilingual Grok 4 leads

Claude Opus 4: 48.8 (#138), Grok 4: 51.8 (#103)

Multilingual benchmarks
BenchmarkClaude Opus 4Grok 4
LMArena Non-English13621403
LMArena Chinese13861427
LMArena French13721418
LMArena German13911429
LMArena Japanese13311394
LMArena Korean13211377
LMArena Russian13921410
LMArena Spanish13891420

Instruction Following Grok 4 leads

Claude Opus 4: 77.1 (#28), Grok 4: 79.2 (#5)

Instruction Following benchmarks
BenchmarkClaude Opus 4Grok 4
IFEval91.8%94.9%
LMArena Instruction Following14061387

Long Context Grok 4 leads

Claude Opus 4: 39.6 (#172), Grok 4: 63.1 (#4)

Long Context benchmarks
BenchmarkClaude Opus 4Grok 4
Fiction.LiveBench61.1%94.4%
LMArena Longer Query14221409

Writing & Preference Claude Opus 4 leads

Claude Opus 4: 61.2 (#89), Grok 4: 58.5 (#116)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Grok 4
LMArena Text13771411
LMArena Creative Writing13871397
Short-Story Creative Writing83.6%76.9%
WildBench85.2%79.7%
LMArena Multi-Turn13961416
EQ-Bench Creative Writing1580—

Frequently asked questions

Is Claude Opus 4 better than Grok 4?

Grok 4 is the stronger model overall, scoring 48.1 to 43.1 on the Noometry Index.

Is Claude Opus 4 or Grok 4 better for coding?

Grok 4 scores higher on coding benchmarks: 50.3 versus 47.2 in the Noometry coding category.

How many benchmarks do Claude Opus 4 and Grok 4 share?

43 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Grok 4 has 48.

Related comparisons

Go deeper