Model comparison

Claude Sonnet 4 vs Gemini 2.5 Pro

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 40.8 on the Noometry Index.

Last verified . 55 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Gemini 2.5 Pro Google

45.0

Rank #75 Confirmed

Summary

  • They share 55 benchmarks with published results for both. Claude Sonnet 4 scores higher in 3 categories and Gemini 2.5 Pro in 7 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Gemini 2.5 Pro leads 59.8 to 33.7.
  • The biggest single-benchmark swing is GeoBench: 37% for Claude Sonnet 4 and 86% for Gemini 2.5 Pro.
  • Gemini 2.5 Pro is cheaper at $1.25 / $10 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • Gemini 2.5 Pro accepts more context: 1.05M tokens versus 200K.

Side by side

Claude Sonnet 4 and Gemini 2.5 Pro specifications
Claude Sonnet 4Gemini 2.5 Pro
ProviderAnthropicGoogle
Noometry Index40.845.0
Released2025-05-222025-03-25
WeightsProprietaryProprietary
Context window200K1.05M
Max output64K66K
Input $ / M tokens$3$1.25
Output $ / M tokens$15$10
Results tracked5878

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude Sonnet 4: 43.5 (#88), Gemini 2.5 Pro: 42.4 (#101)

Coding benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
SWE-bench Verified (bash only)64.9%53.6%
Aider Polyglot61.3%83.1%
SciCode40%42.8%
GSO4.9%3.9%
WeirdML46.1%54%
LMArena Coding14141452
ALE-Bench655.35785.52
SWE-bench Verified—57.6%
LMArena WebDev—1227
LiveBench Coding—85.9%
CadEval—64%
AlgoTune—1.51

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), Gemini 2.5 Pro: 29.2 (#88)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
TheAgentCompany33.1%30.3%
DeepResearch Bench46.6%42.8%
METR Time Horizons62%55.4%
Terminal-Bench—32.6%
GDPval—23.3%
Remote Labor Index—0.8%
τ²-bench Banking—13.7%
Cybench35%—
OSWorld43.9%—
BALROG—43.3%
LMArena Search—1142
Vending-Bench 2—573.64

Reasoning Gemini 2.5 Pro leads

Claude Sonnet 4: 22.9 (#187), Gemini 2.5 Pro: 28.8 (#99)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
ARC-AGI-25.9%4.9%
SimpleBench45.5%62.4%
Kagi LLM Benchmark73%70.3%
ARC-AGI-140%41%
CritPt0.3%2%
EnigmaEval3.1%5.6%
LMArena Hard Prompts13721455
DTBench77.1%82.4%
LMCA29%34.8%
Epoch Capabilities Index141.69145.32
ForecastBench60.261.3
Chess Puzzles—20%
LiveBench Reasoning—89.8%
LiveBench Data Analysis—79.9%
LiveBench—82.3%

Math Claude Sonnet 4 leads

Claude Sonnet 4: 43.3 (#80), Gemini 2.5 Pro: 32.5 (#213)

Math benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
OTIS Mock AIME 2024-202571.1%84.7%
Omni-MATH60.2%41.6%
LMArena Math13751450
MATH Level 584.4%95.9%
FrontierMath (Feb 2025 set)4.1%14.1%
FrontierMath Tier 4 (v1)0%4.2%
FrontierMath (Tiers 1-3)—24.6%
FrontierMath Tier 4—0%
LiveBench Math—90.2%

Knowledge Gemini 2.5 Pro leads

Claude Sonnet 4: 41.8 (#108), Gemini 2.5 Pro: 56.0 (#46)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
GPQA Diamond79.2%85.3%
Humanity's Last Exam7.8%21.6%
MMLU-Pro84.3%86.3%
Confabulations13.2%10.6%
Vectara Hallucination Rate10.3%7%
GPQA (HELM)70.6%74.9%
LMArena Expert13721452

Multimodal Gemini 2.5 Pro leads

Claude Sonnet 4: 26.2 (#121), Gemini 2.5 Pro: 45.2 (#18)

Multimodal benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
LMArena Vision11911263
GeoBench37%86%
VPCT34%48%
LMArena Document—1421
MindCube44.8%—
SpatialViz-Bench—44.7%

Multilingual Gemini 2.5 Pro leads

Claude Sonnet 4: 46.7 (#156), Gemini 2.5 Pro: 55.3 (#31)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
LMArena Non-English13331451
LMArena Chinese13501507
LMArena French13631472
LMArena German13311487
LMArena Japanese13021461
LMArena Korean12911434
LMArena Russian13551461
LMArena Spanish13571473

Instruction Following Gemini 2.5 Pro leads

Claude Sonnet 4: 71.7 (#145), Gemini 2.5 Pro: 75.0 (#75)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
IFEval84%84%
LMArena Instruction Following13761437
LiveBench Instruction Following—80.6%

Long Context Gemini 2.5 Pro leads

Claude Sonnet 4: 33.7 (#259), Gemini 2.5 Pro: 59.8 (#5)

Long Context benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
Fiction.LiveBench46.9%91.7%
LMArena Longer Query13981449

Writing & Preference Gemini 2.5 Pro leads

Claude Sonnet 4: 57.1 (#132), Gemini 2.5 Pro: 63.7 (#62)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Gemini 2.5 Pro
LMArena Text13511458
LMArena Creative Writing13451454
Short-Story Creative Writing81.4%83.8%
EQ-Bench Creative Writing14831421
WildBench83.8%85.7%
LMArena Multi-Turn13761453
LiveBench Language—67.8%

Frequently asked questions

Is Claude Sonnet 4 better than Gemini 2.5 Pro?

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or Gemini 2.5 Pro?

Gemini 2.5 Pro is cheaper. It lists at $1.25 per million input tokens and $10 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or Gemini 2.5 Pro better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Pro does, with 1.05M tokens against 200K.

How many benchmarks do Claude Sonnet 4 and Gemini 2.5 Pro share?

55 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Gemini 2.5 Pro has 78.

Related comparisons

Go deeper