Model comparison

Claude Sonnet 4 vs Gemini 2.0 Flash (Feb 2025)

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 35.1 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Sonnet 4 scores higher in 6 categories and Gemini 2.0 Flash (Feb 2025) in 4 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Claude Sonnet 4 leads 43.5 to 28.4.
  • The biggest single-benchmark swing is SWE-bench Verified (bash only): 64.9% for Claude Sonnet 4 and 13.5% for Gemini 2.0 Flash (Feb 2025).

Side by side

Claude Sonnet 4 and Gemini 2.0 Flash (Feb 2025) specifications
Claude Sonnet 4Gemini 2.0 Flash (Feb 2025)
ProviderAnthropicGoogle
Noometry Index40.835.1
Released2025-05-222024-12-06
WeightsProprietaryProprietary
Context window200K—
Max output64K—
Input $ / M tokens$3—
Output $ / M tokens$15—
Results tracked5854

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude Sonnet 4: 43.5 (#88), Gemini 2.0 Flash (Feb 2025): 28.4 (#315)

Coding benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
SWE-bench Verified (bash only)64.9%13.5%
Aider Polyglot61.3%38.2%
WeirdML46.1%25.8%
LMArena Coding14141350
SciCode40%—
GSO4.9%—
BigCodeBench Instruct—45.9%
LiveBench Coding—63.4%
BigCodeBench Complete—59.9%
CadEval—30%
ALE-Bench655.35—

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), Gemini 2.0 Flash (Feb 2025): 28.1 (#92)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
TheAgentCompany33.1%11.4%
Cybench35%—
DeepResearch Bench46.6%—
OSWorld43.9%—
METR Time Horizons62%—

Reasoning Claude Sonnet 4 leads

Claude Sonnet 4: 22.9 (#187), Gemini 2.0 Flash (Feb 2025): 15.2 (#318)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
ARC-AGI-25.9%1.3%
SimpleBench45.5%31.1%
Kagi LLM Benchmark73%37.8%
EnigmaEval3.1%1.1%
LMArena Hard Prompts13721346
DTBench77.1%63.2%
Epoch Capabilities Index141.69135.36
ARC-AGI-140%—
CritPt0.3%—
LiveBench Reasoning—78.2%
LiveBench Data Analysis—69.4%
LMCA29%—
ForecastBench60.2—
LiveBench—66.9%

Math Claude Sonnet 4 leads

Claude Sonnet 4: 43.3 (#80), Gemini 2.0 Flash (Feb 2025): 37.9 (#146)

Math benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
OTIS Mock AIME 2024-202571.1%57.8%
Omni-MATH60.2%45.9%
LMArena Math13751352
MATH Level 584.4%82.2%
FrontierMath (Feb 2025 set)4.1%1.7%
LiveBench Math—75.8%
FrontierMath Tier 4 (v1)0%—

Knowledge Claude Sonnet 4 leads

Claude Sonnet 4: 41.8 (#108), Gemini 2.0 Flash (Feb 2025): 32.0 (#213)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
GPQA Diamond79.2%64.1%
Humanity's Last Exam7.8%6.6%
MMLU-Pro84.3%73.7%
Confabulations13.2%12.4%
GPQA (HELM)70.6%55.6%
LMArena Expert13721339
Vectara Hallucination Rate10.3%—
MMLU—79.7%

Multimodal Gemini 2.0 Flash (Feb 2025) leads

Claude Sonnet 4: 26.2 (#121), Gemini 2.0 Flash (Feb 2025): 36.5 (#79)

Multimodal benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
LMArena Vision11911158
GeoBench37%77%
VPCT34%—
MindCube44.8%—

Multilingual Too close to call

Claude Sonnet 4: 46.7 (#156), Gemini 2.0 Flash (Feb 2025): 47.4 (#149)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
LMArena Non-English13331342
LMArena Chinese13501373
LMArena French13631391
LMArena German13311353
LMArena Japanese13021294
LMArena Korean12911313
LMArena Russian13551351
LMArena Spanish13571363

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Claude Sonnet 4: 71.7 (#145), Gemini 2.0 Flash (Feb 2025): 74.4 (#97)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
IFEval84%84.1%
LMArena Instruction Following13761336
LiveBench Instruction Following—85.8%

Long Context Gemini 2.0 Flash (Feb 2025) leads

Claude Sonnet 4: 33.7 (#259), Gemini 2.0 Flash (Feb 2025): 38.1 (#203)

Long Context benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
Fiction.LiveBench46.9%61.1%
LMArena Longer Query13981344

Writing & Preference Claude Sonnet 4 leads

Claude Sonnet 4: 57.1 (#132), Gemini 2.0 Flash (Feb 2025): 49.5 (#190)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Gemini 2.0 Flash (Feb 2025)
LMArena Text13511354
LMArena Creative Writing13451340
Short-Story Creative Writing81.4%73.8%
EQ-Bench Creative Writing14831128
WildBench83.8%80%
LMArena Multi-Turn13761350
LiveBench Language—51.3%

Frequently asked questions

Is Claude Sonnet 4 better than Gemini 2.0 Flash (Feb 2025)?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 35.1 on the Noometry Index.

Is Claude Sonnet 4 or Gemini 2.0 Flash (Feb 2025) better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 28.4 in the Noometry coding category.

How many benchmarks do Claude Sonnet 4 and Gemini 2.0 Flash (Feb 2025) share?

43 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Gemini 2.0 Flash (Feb 2025) has 54.

Related comparisons

Go deeper