Model comparison

Claude Sonnet 4 vs Grok 4 Fast

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 39.4 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Grok 4 Fast xAI

39.4

Rank #167 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude Sonnet 4 scores higher in 5 categories and Grok 4 Fast in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok 4 Fast leads 63.2 to 33.7.
  • The biggest single-benchmark swing is Fiction.LiveBench: 46.9% for Claude Sonnet 4 and 94.4% for Grok 4 Fast.

Side by side

Claude Sonnet 4 and Grok 4 Fast specifications
Claude Sonnet 4Grok 4 Fast
ProviderAnthropicxAI
Noometry Index40.839.4
Released2025-05-222025-09-19
WeightsProprietaryProprietary
Context window200K—
Max output64K—
Input $ / M tokens$3—
Output $ / M tokens$15—
Results tracked5830

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude Sonnet 4: 43.5 (#88), Grok 4 Fast: 32.5 (#271)

Coding benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
WeirdML46.1%42.9%
LMArena Coding14141429
SWE-bench Verified (bash only)64.9%—
Aider Polyglot61.3%—
LMArena WebDev—1159
SciCode40%—
GSO4.9%—
ALE-Bench655.35—

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), Grok 4 Fast: 29.5 (#86)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
Cybench35%30%
TheAgentCompany33.1%—
τ²-bench Banking—15.7%
DeepResearch Bench46.6%—
OSWorld43.9%—
LMArena Search—1171
METR Time Horizons62%—

Reasoning Too close to call

Claude Sonnet 4: 22.9 (#187), Grok 4 Fast: 22.2 (#201)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
ARC-AGI-25.9%5.3%
Kagi LLM Benchmark73%66.1%
ARC-AGI-140%48.5%
LMArena Hard Prompts13721412
DTBench77.1%82.7%
Epoch Capabilities Index141.69144.2
ForecastBench60.260.5
SimpleBench45.5%—
CritPt0.3%—
EnigmaEval3.1%—
LMCA29%—

Math Claude Sonnet 4 leads

Claude Sonnet 4: 43.3 (#80), Grok 4 Fast: 38.9 (#123)

Math benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
LMArena Math13751419
OTIS Mock AIME 2024-202571.1%—
Omni-MATH60.2%—
MATH Level 584.4%—
FrontierMath (Feb 2025 set)4.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Claude Sonnet 4 leads

Claude Sonnet 4: 41.8 (#108), Grok 4 Fast: 32.0 (#214)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
Vectara Hallucination Rate10.3%19.7%
LMArena Expert13721411
GPQA Diamond79.2%—
Humanity's Last Exam7.8%—
MMLU-Pro84.3%—
Confabulations13.2%—
GPQA (HELM)70.6%—

Multimodal Not comparable

Claude Sonnet 4: 26.2 (#121), Grok 4 Fast: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
LMArena Vision1191—
GeoBench37%—
VPCT34%—
MindCube44.8%—

Multilingual Grok 4 Fast leads

Claude Sonnet 4: 46.7 (#156), Grok 4 Fast: 51.3 (#111)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
LMArena Non-English13331396
LMArena Chinese13501457
LMArena French13631431
LMArena German13311383
LMArena Japanese13021352
LMArena Korean12911357
LMArena Russian13551389
LMArena Spanish13571415

Instruction Following Grok 4 Fast leads

Claude Sonnet 4: 71.7 (#145), Grok 4 Fast: 73.2 (#121)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
LMArena Instruction Following13761387
IFEval84%—

Long Context Grok 4 Fast leads

Claude Sonnet 4: 33.7 (#259), Grok 4 Fast: 63.2 (#3)

Long Context benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
Fiction.LiveBench46.9%94.4%
LMArena Longer Query13981415

Writing & Preference Grok 4 Fast leads

Claude Sonnet 4: 57.1 (#132), Grok 4 Fast: 60.0 (#102)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Grok 4 Fast
LMArena Text13511407
LMArena Creative Writing13451387
LMArena Multi-Turn13761414
Short-Story Creative Writing81.4%—
EQ-Bench Creative Writing1483—
WildBench83.8%—

Frequently asked questions

Is Claude Sonnet 4 better than Grok 4 Fast?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 39.4 on the Noometry Index.

Is Claude Sonnet 4 or Grok 4 Fast better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 32.5 in the Noometry coding category.

How many benchmarks do Claude Sonnet 4 and Grok 4 Fast share?

27 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Grok 4 Fast has 30.

Related comparisons

Go deeper