Model comparison

Claude 3.7 Sonnet vs Gemini 1.5 Flash (May 2024)

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 33.2 on the Noometry Index.

Last verified . 30 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 8 categories and Gemini 1.5 Flash (May 2024) in 2 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude 3.7 Sonnet leads 37.5 to 22.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Claude 3.7 Sonnet and 16.3% for Gemini 1.5 Flash (May 2024).

Side by side

Claude 3.7 Sonnet and Gemini 1.5 Flash (May 2024) specifications
Claude 3.7 SonnetGemini 1.5 Flash (May 2024)
ProviderAnthropicGoogle
Noometry Index39.533.2
Released2025-02-242024-05-14
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked5842

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
LMArena Coding13611261
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
GSO3.8%—
WeirdML—24.9%
BigCodeBench Instruct—43.5%
LiveBench Coding74.5%—
BigCodeBench Complete—55.1%
CadEval54%—
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 34.1 (#50), Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
BALROG—14.6%
METR Time Horizons60%—

Reasoning Gemini 1.5 Flash (May 2024) leads

Claude 3.7 Sonnet: 18.6 (#277), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
LMArena Hard Prompts13331257
Epoch Capabilities Index141.16129.36
ForecastBench61.853.9
ARC-AGI-20.9%—
SimpleBench46.4%—
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
DTBench—53.8%
LiveBench Data Analysis74%—
LiveBench76.1%—
PIQA—87.5%

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
OTIS Mock AIME 2024-202557.8%16.3%
Omni-MATH33%30.4%
LMArena Math13371269
MATH Level 591.2%61.9%
FrontierMath (Feb 2025 set)4.1%0%
LiveBench Math79%—
GSM8K—82.4%

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
GPQA Diamond79.7%47.3%
MMLU-Pro78.4%67.8%
GPQA (HELM)60.8%43.7%
LMArena Expert13211233
Humanity's Last Exam8%—
Confabulations14.7%—
BoolQ—85.8%
MMLU—77.9%

Multimodal Gemini 1.5 Flash (May 2024) leads

Claude 3.7 Sonnet: 33.7 (#95), Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
LMArena Vision11691141
GeoBench68%76%
Video-MME—70.3%
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 44.1 (#179), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
LMArena Non-English12961278
LMArena Chinese12991295
LMArena French13031258
LMArena German13011262
LMArena Japanese12671252
LMArena Korean12491221
LMArena Russian13111288
LMArena Spanish12981243

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
IFEval83.4%83.1%
LMArena Instruction Following13521258
LiveBench Instruction Following81.3%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
LMArena Longer Query13731284
Fiction.LiveBench83.3%—

Writing & Preference Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 54.4 (#150), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetGemini 1.5 Flash (May 2024)
LMArena Text13141287
LMArena Creative Writing13321285
WildBench81.4%79.2%
LMArena Multi-Turn13391253
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Gemini 1.5 Flash (May 2024)?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 33.2 on the Noometry Index.

Is Claude 3.7 Sonnet or Gemini 1.5 Flash (May 2024) better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Gemini 1.5 Flash (May 2024) share?

30 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper