Model comparison

Claude 3.5 Sonnet vs Gemini 1.5 Pro (May 2024)

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 32.1 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 6 categories and Gemini 1.5 Pro (May 2024) in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude 3.5 Sonnet leads 32.3 to 17.9.
  • The biggest single-benchmark swing is TheAgentCompany: 24% for Claude 3.5 Sonnet and 3.4% for Gemini 1.5 Pro (May 2024).

Side by side

Claude 3.5 Sonnet and Gemini 1.5 Pro (May 2024) specifications
Claude 3.5 SonnetGemini 1.5 Pro (May 2024)
ProviderAnthropicGoogle
Noometry Index34.632.1
Released2024-06-202024-02-15
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6045

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Gemini 1.5 Pro (May 2024): 34.2 (#241)

Coding benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
WeirdML40%22.2%
BigCodeBench Instruct46.8%43.8%
LMArena Coding13421294
BigCodeBench Complete58.6%57.5%
CadEval48%34%
HumanEval+81.7%79.3%
MBPP+74.3%74.6%
Aider Polyglot51.6%—
GSO4.6%—
LiveBench Coding67.1%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), Gemini 1.5 Pro (May 2024): 17.9 (#145)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
TheAgentCompany24%3.4%
Cybench17.5%7.5%
BALROG32.6%21%
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Gemini 1.5 Pro (May 2024): 12.3 (#338)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
SimpleBench41.4%27.1%
LMArena Hard Prompts13051296
DTBench67.8%59%
Epoch Capabilities Index133.55131.73
ForecastBench60.758.4
ARC-AGI-2—0.8%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LiveBench Data Analysis55%—
BIG-Bench Hard—89.2%
LiveBench59%—

Math Gemini 1.5 Pro (May 2024) leads

Claude 3.5 Sonnet: 19.2 (#288), Gemini 1.5 Pro (May 2024): 25.8 (#266)

Math benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
OTIS Mock AIME 2024-20258.5%23.1%
Omni-MATH27.6%36.4%
LMArena Math13071315
MATH Level 556.9%70.4%
LiveBench Math52.3%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Too close to call

Claude 3.5 Sonnet: 28.6 (#245), Gemini 1.5 Pro (May 2024): 29.4 (#239)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
GPQA Diamond55.3%57.2%
Humanity's Last Exam4.1%4.6%
MMLU-Pro77.7%73.7%
Confabulations19.9%13.5%
GPQA (HELM)56.5%53.4%
LMArena Expert12651279
MMLU87.3%86.9%

Multimodal Gemini 1.5 Pro (May 2024) leads

Claude 3.5 Sonnet: 26.5 (#120), Gemini 1.5 Pro (May 2024): 36.8 (#77)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
LMArena Vision11251161
Video-MME60%75%
GeoBench62%—
VPCT33%—

Multilingual Gemini 1.5 Pro (May 2024) leads

Claude 3.5 Sonnet: 43.2 (#185), Gemini 1.5 Pro (May 2024): 45.3 (#174)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
LMArena Non-English12831312
LMArena Chinese12721331
LMArena French13051302
LMArena German12971286
LMArena Japanese12341292
LMArena Korean12001298
LMArena Russian13061320
LMArena Spanish12901311

Instruction Following Too close to call

Claude 3.5 Sonnet: 68.8 (#182), Gemini 1.5 Pro (May 2024): 68.6 (#185)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
IFEval85.5%83.7%
LMArena Instruction Following12971297
LiveBench Instruction Following69.3%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Gemini 1.5 Pro (May 2024): 39.8 (#169)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
LMArena Longer Query13111308

Writing & Preference Too close to call

Claude 3.5 Sonnet: 52.9 (#164), Gemini 1.5 Pro (May 2024): 52.4 (#172)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetGemini 1.5 Pro (May 2024)
LMArena Text12981319
LMArena Creative Writing12921333
WildBench79.2%81.3%
LMArena Multi-Turn13261296
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Gemini 1.5 Pro (May 2024)?

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 32.1 on the Noometry Index.

Is Claude 3.5 Sonnet or Gemini 1.5 Pro (May 2024) better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 34.2 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Gemini 1.5 Pro (May 2024) share?

43 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Gemini 1.5 Pro (May 2024) has 45.

Related comparisons

Go deeper