Model comparison

Claude Sonnet 4 vs Gemini 1.5 Pro (May 2024)

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 32.1 on the Noometry Index.

Last verified . 36 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Summary

  • They share 36 benchmarks with published results for both. Claude Sonnet 4 scores higher in 8 categories and Gemini 1.5 Pro (May 2024) in 2 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude Sonnet 4 leads 38.5 to 17.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 71.1% for Claude Sonnet 4 and 23.1% for Gemini 1.5 Pro (May 2024).

Side by side

Claude Sonnet 4 and Gemini 1.5 Pro (May 2024) specifications
Claude Sonnet 4Gemini 1.5 Pro (May 2024)
ProviderAnthropicGoogle
Noometry Index40.832.1
Released2025-05-222024-02-15
WeightsProprietaryProprietary
Context window200K—
Max output64K—
Input $ / M tokens$3—
Output $ / M tokens$15—
Results tracked5845

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude Sonnet 4: 43.5 (#88), Gemini 1.5 Pro (May 2024): 34.2 (#241)

Coding benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
WeirdML46.1%22.2%
LMArena Coding14141294
SWE-bench Verified (bash only)64.9%—
Aider Polyglot61.3%—
SciCode40%—
GSO4.9%—
BigCodeBench Instruct—43.8%
BigCodeBench Complete—57.5%
CadEval—34%
ALE-Bench655.35—
HumanEval+—79.3%
MBPP+—74.6%

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), Gemini 1.5 Pro (May 2024): 17.9 (#145)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
TheAgentCompany33.1%3.4%
Cybench35%7.5%
DeepResearch Bench46.6%—
OSWorld43.9%—
BALROG—21%
METR Time Horizons62%—

Reasoning Claude Sonnet 4 leads

Claude Sonnet 4: 22.9 (#187), Gemini 1.5 Pro (May 2024): 12.3 (#338)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
ARC-AGI-25.9%0.8%
SimpleBench45.5%27.1%
LMArena Hard Prompts13721296
DTBench77.1%59%
Epoch Capabilities Index141.69131.73
ForecastBench60.258.4
Kagi LLM Benchmark73%—
ARC-AGI-140%—
CritPt0.3%—
EnigmaEval3.1%—
LMCA29%—
BIG-Bench Hard—89.2%

Math Claude Sonnet 4 leads

Claude Sonnet 4: 43.3 (#80), Gemini 1.5 Pro (May 2024): 25.8 (#266)

Math benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
OTIS Mock AIME 2024-202571.1%23.1%
Omni-MATH60.2%36.4%
LMArena Math13751315
MATH Level 584.4%70.4%
FrontierMath (Feb 2025 set)4.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Claude Sonnet 4 leads

Claude Sonnet 4: 41.8 (#108), Gemini 1.5 Pro (May 2024): 29.4 (#239)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
GPQA Diamond79.2%57.2%
Humanity's Last Exam7.8%4.6%
MMLU-Pro84.3%73.7%
Confabulations13.2%13.5%
GPQA (HELM)70.6%53.4%
LMArena Expert13721279
Vectara Hallucination Rate10.3%—
MMLU—86.9%

Multimodal Gemini 1.5 Pro (May 2024) leads

Claude Sonnet 4: 26.2 (#121), Gemini 1.5 Pro (May 2024): 36.8 (#77)

Multimodal benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
LMArena Vision11911161
Video-MME—75%
GeoBench37%—
VPCT34%—
MindCube44.8%—

Multilingual Claude Sonnet 4 leads

Claude Sonnet 4: 46.7 (#156), Gemini 1.5 Pro (May 2024): 45.3 (#174)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
LMArena Non-English13331312
LMArena Chinese13501331
LMArena French13631302
LMArena German13311286
LMArena Japanese13021292
LMArena Korean12911298
LMArena Russian13551320
LMArena Spanish13571311

Instruction Following Claude Sonnet 4 leads

Claude Sonnet 4: 71.7 (#145), Gemini 1.5 Pro (May 2024): 68.6 (#185)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
IFEval84%83.7%
LMArena Instruction Following13761297

Long Context Gemini 1.5 Pro (May 2024) leads

Claude Sonnet 4: 33.7 (#259), Gemini 1.5 Pro (May 2024): 39.8 (#169)

Long Context benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
LMArena Longer Query13981308
Fiction.LiveBench46.9%—

Writing & Preference Claude Sonnet 4 leads

Claude Sonnet 4: 57.1 (#132), Gemini 1.5 Pro (May 2024): 52.4 (#172)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Gemini 1.5 Pro (May 2024)
LMArena Text13511319
LMArena Creative Writing13451333
WildBench83.8%81.3%
LMArena Multi-Turn13761296
Short-Story Creative Writing81.4%—
EQ-Bench Creative Writing1483—

Frequently asked questions

Is Claude Sonnet 4 better than Gemini 1.5 Pro (May 2024)?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 32.1 on the Noometry Index.

Is Claude Sonnet 4 or Gemini 1.5 Pro (May 2024) better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 34.2 in the Noometry coding category.

How many benchmarks do Claude Sonnet 4 and Gemini 1.5 Pro (May 2024) share?

36 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Gemini 1.5 Pro (May 2024) has 45.

Related comparisons

Go deeper