Model comparison

Claude Haiku 4.5 vs Gemini 1.5 Flash (May 2024)

Claude Haiku 4.5 is the stronger model overall, scoring 39.5 to 33.2 on the Noometry Index.

Last verified . 31 shared benchmarks.

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Claude Haiku 4.5 scores higher in 8 categories and Gemini 1.5 Flash (May 2024) in 2 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Haiku 4.5 leads 44.9 to 22.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 66.7% for Claude Haiku 4.5 and 16.3% for Gemini 1.5 Flash (May 2024).

Side by side

Claude Haiku 4.5 and Gemini 1.5 Flash (May 2024) specifications
Claude Haiku 4.5Gemini 1.5 Flash (May 2024)
ProviderAnthropicGoogle
Noometry Index39.533.2
Released2025-10-152024-05-14
WeightsProprietaryProprietary
Context window200K—
Max output64K—
Input $ / M tokens$1—
Output $ / M tokens$5—
Results tracked5342

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.0 (#78), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
WeirdML45.4%24.9%
LMArena Coding14531261
SWE-bench Verified (bash only)66.6%—
LMArena WebDev1330—
SWE-bench Multilingual64.7%—
SciCode43.3%—
BigCodeBench Instruct—43.5%
BigCodeBench Complete—55.1%
ALE-Bench653.48—
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Claude Haiku 4.5 leads

Claude Haiku 4.5: 33.6 (#52), Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
BALROG31.2%14.6%
Terminal-Bench35.5%—
Berkeley Function Calling Leaderboard68.7%—
DeepResearch Bench45.5%—
ExploitBench13.7%—
Vending-Bench 2458.89—

Reasoning Gemini 1.5 Flash (May 2024) leads

Claude Haiku 4.5: 15.1 (#320), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
LMArena Hard Prompts14201257
DTBench73.6%53.8%
Epoch Capabilities Index142.41129.36
ForecastBench61.453.9
ARC-AGI-24%—
NYT Connections (extended)14.3%—
ARC-AGI-147.7%—
CritPt0%—
Chess Puzzles8%—
LMCA30.9%—
PIQA—87.5%

Math Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.9 (#78), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
OTIS Mock AIME 2024-202566.7%16.3%
Omni-MATH56.1%30.4%
LMArena Math13961269
MATH Level 596.4%61.9%
FrontierMath (Feb 2025 set)5.9%0%
FrontierMath Tier 4 (v1)2.1%—
GSM8K—82.4%

Knowledge Claude Haiku 4.5 leads

Claude Haiku 4.5: 37.7 (#153), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
GPQA Diamond71.2%47.3%
MMLU-Pro77.7%67.8%
GPQA (HELM)60.5%43.7%
LMArena Expert14421233
SimpleQA Verified13.2%—
Vectara Hallucination Rate9.8%—
BoolQ—85.8%
MMLU—77.9%

Multimodal Gemini 1.5 Flash (May 2024) leads

Claude Haiku 4.5: 26.8 (#118), Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
LMArena Vision—1141
Video-MME—70.3%
GeoBench—76%
Blueprint-Bench 20%—
LMArena Document1420—

Multilingual Claude Haiku 4.5 leads

Claude Haiku 4.5: 49.9 (#129), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
LMArena Non-English13771278
LMArena Chinese14171295
LMArena French14081258
LMArena German13751262
LMArena Japanese13391252
LMArena Korean13471221
LMArena Russian13811288
LMArena Spanish14201243

Instruction Following Claude Haiku 4.5 leads

Claude Haiku 4.5: 71.4 (#149), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
IFEval80.1%83.1%
LMArena Instruction Following14141258

Long Context Claude Haiku 4.5 leads

Claude Haiku 4.5: 43.6 (#92), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
LMArena Longer Query14271284

Writing & Preference Claude Haiku 4.5 leads

Claude Haiku 4.5: 57.9 (#123), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkClaude Haiku 4.5Gemini 1.5 Flash (May 2024)
LMArena Text13961287
LMArena Creative Writing13721285
WildBench83.9%79.2%
LMArena Multi-Turn14091253
EQ-Bench 41064—

Frequently asked questions

Is Claude Haiku 4.5 better than Gemini 1.5 Flash (May 2024)?

Claude Haiku 4.5 is the stronger model overall, scoring 39.5 to 33.2 on the Noometry Index.

Is Claude Haiku 4.5 or Gemini 1.5 Flash (May 2024) better for coding?

Claude Haiku 4.5 scores higher on coding benchmarks: 44.0 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude Haiku 4.5 and Gemini 1.5 Flash (May 2024) share?

31 benchmarks have published results for both models. Claude Haiku 4.5 has 53 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper