Model comparison

Claude 3.5 Haiku vs Gemma 7B

Claude 3.5 Haiku and Gemma 7B score almost the same on the Noometry Index (29.2 vs 30.0), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Gemma 7B Google

30.0

Rank #299 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 5 categories and Gemma 7B in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 7B leads 31.2 to 14.7.
  • Gemma 7B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Gemma 7B specifications
Claude 3.5 HaikuGemma 7B
ProviderAnthropicGoogle
Noometry Index29.230.0
Released2024-10-222024-02-21
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Haiku leads

Claude 3.5 Haiku: 32.9 (#265), Gemma 7B: 30.5 (#294)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Coding12861048
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—
HumanEval+—28.7%
MBPP+—43.4%

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Gemma 7B: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
BALROG19.3%—

Reasoning Gemma 7B leads

Claude 3.5 Haiku: 17.7 (#290), Gemma 7B: 19.9 (#249)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Hard Prompts12511042
Epoch Capabilities Index127.15111.99
CritPt0%—
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Adversarial NLI—48.7%
BIG-Bench Hard—55.1%
HellaSwag—82.2%
LiveBench43.5%—
PIQA—81.2%
WinoGrande—79%

Math Gemma 7B leads

Claude 3.5 Haiku: 14.7 (#300), Gemma 7B: 31.2 (#228)

Math benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Math12441066
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—
GSM8K—46.4%

Knowledge Gemma 7B leads

Claude 3.5 Haiku: 18.7 (#281), Gemma 7B: 27.3 (#252)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Expert12081001
MMLU74.3%66.1%
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
ARC (AI2) Challenge—78.3%
BoolQ—83.2%
OpenBookQA—78.6%
TriviaQA—72.3%

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Gemma 7B: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Vision1092—
GeoBench34%—

Multilingual Claude 3.5 Haiku leads

Claude 3.5 Haiku: 40.0 (#218), Gemma 7B: 25.1 (#287)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Non-English1238999
LMArena Chinese12291035
LMArena French12641025
LMArena Russian1253993
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Spanish1261—

Instruction Following Claude 3.5 Haiku leads

Claude 3.5 Haiku: 62.9 (#234), Gemma 7B: 51.5 (#295)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Instruction Following12411017
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Claude 3.5 Haiku leads

Claude 3.5 Haiku: 38.3 (#200), Gemma 7B: 31.1 (#282)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Longer Query12611022

Writing & Preference Claude 3.5 Haiku leads

Claude 3.5 Haiku: 42.7 (#234), Gemma 7B: 27.1 (#302)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGemma 7B
LMArena Text12551056
LMArena Creative Writing12331024
LMArena Multi-Turn1265963
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Gemma 7B?

Claude 3.5 Haiku and Gemma 7B score almost the same on the Noometry Index (29.2 vs 30.0), so choose on price, context window or the category you care about most.

Is Claude 3.5 Haiku or Gemma 7B better for coding?

Claude 3.5 Haiku scores higher on coding benchmarks: 32.9 versus 30.5 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Gemma 7B share?

15 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Gemma 7B has 27.

Related comparisons

Go deeper