Model comparison

Claude 3.5 Haiku vs Gemini 3 Pro

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 29.2 on the Noometry Index.

Last verified . 32 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Gemini 3 Pro in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemini 3 Pro leads 64.4 to 18.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 91.4% for Gemini 3 Pro.

Side by side

Claude 3.5 Haiku and Gemini 3 Pro specifications
Claude 3.5 HaikuGemini 3 Pro
ProviderAnthropicGoogle
Noometry Index29.254.8
Released2024-10-222025-11-18
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4967

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3 Pro leads

Claude 3.5 Haiku: 32.9 (#265), Gemini 3 Pro: 51.6 (#39)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
WeirdML30.7%69.9%
LMArena Coding12861481
SWE-bench Verified—72.9%
SWE-bench Verified (bash only)—74.2%
Aider Polyglot28%—
LMArena WebDev—1440
SWE-bench Multilingual—68.7%
SciCode27.4%—
GSO—18.6%
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—
ALE-Bench—1,177
AlgoTune—1.83

Agentic & Tool Use Gemini 3 Pro leads

Claude 3.5 Haiku: 28.0 (#95), Gemini 3 Pro: 40.6 (#23)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
BALROG19.3%58.1%
Terminal-Bench—69.4%
Berkeley Function Calling Leaderboard—72.5%
GDPval—40.3%
Remote Labor Index—1.3%
τ²-bench Airline—80.5%
τ²-bench Banking—18%
τ²-bench Retail—75.9%
τ²-bench Telecom—91%
DeepResearch Bench—46.3%
LMArena Search—1207
METR Time Horizons—71%
Vending-Bench 2—5,478

Reasoning Gemini 3 Pro leads

Claude 3.5 Haiku: 17.7 (#290), Gemini 3 Pro: 52.5 (#31)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
CritPt0%6.9%
LMArena Hard Prompts12511480
Epoch Capabilities Index127.15152.92
ARC-AGI-2—31.1%
SimpleBench—76.4%
Kagi LLM Benchmark—80.1%
NYT Connections (extended)—94.4%
ARC-AGI-1—75%
Chess Puzzles—31%
EnigmaEval—18.2%
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
ForecastBench—61.2
LiveBench43.5%—

Math Gemini 3 Pro leads

Claude 3.5 Haiku: 14.7 (#300), Gemini 3 Pro: 49.9 (#59)

Math benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
OTIS Mock AIME 2024-20254.3%91.4%
Omni-MATH22.4%55.5%
LMArena Math12441476
FrontierMath (Feb 2025 set)0.3%37.6%
MathArena Final-Answer Competitions—67%
ProofBench—20%
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath Tier 4 (v1)—18.8%

Knowledge Gemini 3 Pro leads

Claude 3.5 Haiku: 18.7 (#281), Gemini 3 Pro: 64.4 (#16)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
GPQA Diamond38.1%92.6%
MMLU-Pro60.5%90.3%
GPQA (HELM)36.3%80.3%
LMArena Expert12081475
Humanity's Last Exam—37.5%
Confabulations36.7%—
Vectara Hallucination Rate—13.6%
MMLU74.3%—

Multimodal Gemini 3 Pro leads

Claude 3.5 Haiku: 26.8 (#117), Gemini 3 Pro: 57.6 (#2)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
LMArena Vision10921305
GeoBench34%84%
VPCT—91%
LMArena Document—1434

Multilingual Gemini 3 Pro leads

Claude 3.5 Haiku: 40.0 (#218), Gemini 3 Pro: 56.9 (#16)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
LMArena Non-English12381474
LMArena Chinese12291523
LMArena French12641492
LMArena German12371515
LMArena Japanese11751510
LMArena Korean11731448
LMArena Russian12531493
LMArena Spanish12611470

Instruction Following Gemini 3 Pro leads

Claude 3.5 Haiku: 62.9 (#234), Gemini 3 Pro: 76.3 (#45)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
IFEval79.2%87.7%
LMArena Instruction Following12411458
LiveBench Instruction Following61.9%—

Long Context Gemini 3 Pro leads

Claude 3.5 Haiku: 38.3 (#200), Gemini 3 Pro: 44.0 (#79)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
LMArena Longer Query12611471
CL-bench—15.8%

Writing & Preference Gemini 3 Pro leads

Claude 3.5 Haiku: 42.7 (#234), Gemini 3 Pro: 66.4 (#35)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGemini 3 Pro
LMArena Text12551479
LMArena Creative Writing12331482
EQ-Bench Creative Writing11461525
WildBench76%85.9%
LMArena Multi-Turn12651484
Short-Story Creative Writing73.5%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Gemini 3 Pro?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Gemini 3 Pro better for coding?

Gemini 3 Pro scores higher on coding benchmarks: 51.6 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Gemini 3 Pro share?

32 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Gemini 3 Pro has 67.

Related comparisons

Go deeper