Model comparison

Claude 3 Haiku vs Gemma 2 9B

Claude 3 Haiku and Gemma 2 9B score almost the same on the Noometry Index (25.9 vs 25.9), so choose on price, context window or the category you care about most.

Last verified . 25 shared benchmarks.

Claude 3 Haiku Anthropic

25.9

Rank #340 Confirmed

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3 Haiku scores higher in 3 categories and Gemma 2 9B in 5 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Claude 3 Haiku leads 17.3 to 9.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 50.1% for Claude 3 Haiku and 40.6% for Gemma 2 9B.
  • Gemma 2 9B has downloadable open weights; the other is API-only.

Side by side

Claude 3 Haiku and Gemma 2 9B specifications
Claude 3 HaikuGemma 2 9B
ProviderAnthropicGoogle
Noometry Index25.925.9
Released2024-03-072024-06-24
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3735

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 9B leads

Claude 3 Haiku: 26.4 (#325), Gemma 2 9B: 29.4 (#304)

Coding benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
BigCodeBench Instruct39.4%34.7%
LMArena Coding11991173
BigCodeBench Complete50.1%40.6%
WeirdML9.8%—
LiveBench Coding—22.5%
CadEval12%—
HumanEval+68.9%—
MBPP+68.8%—

Reasoning Too close to call

Claude 3 Haiku: 16.3 (#307), Gemma 2 9B: 15.9 (#309)

Reasoning benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
LMArena Hard Prompts11741171
Epoch Capabilities Index118.35119.83
Kagi LLM Benchmark34.2%—
LiveBench Reasoning—15.2%
DTBench50.1%—
LiveBench Data Analysis—36.4%
LMCA8.8%—
ForecastBench53.2—
LiveBench—28.7%
PIQA—83.7%
WinoGrande74.2%—

Math Too close to call

Claude 3 Haiku: 9.8 (#319), Gemma 2 9B: 9.9 (#318)

Math benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
OTIS Mock AIME 2024-20251.8%0.6%
LMArena Math11881183
MATH Level 514.9%21%
LiveBench Math—19.8%
GSM8K—84.9%

Knowledge Claude 3 Haiku leads

Claude 3 Haiku: 17.3 (#285), Gemma 2 9B: 9.7 (#305)

Knowledge benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
GPQA Diamond36.3%27.5%
LMArena Expert11481147
MMLU73.8%72.1%
Confabulations34.2%—
BoolQ—85.7%

Multimodal Not comparable

Claude 3 Haiku: 23.6 (#128), Gemma 2 9B: —

Multimodal benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
LMArena Vision950—
ScienceQA72%—

Multilingual Too close to call

Claude 3 Haiku: 36.0 (#243), Gemma 2 9B: 36.6 (#238)

Multilingual benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
LMArena Non-English11781188
LMArena Chinese11551185
LMArena French11951190
LMArena German11741186
LMArena Japanese11021144
LMArena Korean11091137
LMArena Russian12041200
LMArena Spanish11661200

Instruction Following Claude 3 Haiku leads

Claude 3 Haiku: 61.3 (#247), Gemma 2 9B: 57.6 (#269)

Instruction Following benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
LMArena Instruction Following11731178
LiveBench Instruction Following—52.6%

Long Context Too close to call

Claude 3 Haiku: 36.1 (#237), Gemma 2 9B: 36.3 (#233)

Long Context benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
LMArena Longer Query11901197

Writing & Preference Gemma 2 9B leads

Claude 3 Haiku: 29.7 (#291), Gemma 2 9B: 32.1 (#281)

Writing & Preference benchmarks
BenchmarkClaude 3 HaikuGemma 2 9B
LMArena Text11951207
LMArena Creative Writing11571206
EQ-Bench Creative Writing717841
LMArena Multi-Turn11901193
LiveBench Language—25.5%

Frequently asked questions

Is Claude 3 Haiku better than Gemma 2 9B?

Claude 3 Haiku and Gemma 2 9B score almost the same on the Noometry Index (25.9 vs 25.9), so choose on price, context window or the category you care about most.

Is Claude 3 Haiku or Gemma 2 9B better for coding?

Gemma 2 9B scores higher on coding benchmarks: 29.4 versus 26.4 in the Noometry coding category.

How many benchmarks do Claude 3 Haiku and Gemma 2 9B share?

25 benchmarks have published results for both models. Claude 3 Haiku has 37 scored results on Noometry and Gemma 2 9B has 35.

Related comparisons

Go deeper