Model comparison

Claude 2.1 vs Gemma 7B

Gemma 7B is the stronger model overall, scoring 30.0 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Gemma 7B Google

30.0

Rank #299 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 1 category and Gemma 7B in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 7B leads 31.2 to 10.2.
  • Gemma 7B has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Gemma 7B specifications
Claude 2.1Gemma 7B
ProviderAnthropicGoogle
Noometry Index25.230.0
Released2023-11-212024-02-21
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked727

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 7B leads

Claude 2.1: 26.2 (#327), Gemma 7B: 30.5 (#294)

Coding benchmarks
BenchmarkClaude 2.1Gemma 7B
WeirdML7.1%—
LMArena Coding—1048
HumanEval+—28.7%
MBPP+—43.4%

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Gemma 7B: 19.9 (#249)

Reasoning benchmarks
BenchmarkClaude 2.1Gemma 7B
Epoch Capabilities Index119.27111.99
LMArena Hard Prompts—1042
DTBench51%—
Adversarial NLI—48.7%
BIG-Bench Hard—55.1%
ForecastBench54.2—
HellaSwag—82.2%
PIQA—81.2%
WinoGrande—79%

Math Gemma 7B leads

Claude 2.1: 10.2 (#315), Gemma 7B: 31.2 (#228)

Math benchmarks
BenchmarkClaude 2.1Gemma 7B
OTIS Mock AIME 2024-20251.9%—
LMArena Math—1066
GSM8K—46.4%

Knowledge Gemma 7B leads

Claude 2.1: 15.4 (#292), Gemma 7B: 27.3 (#252)

Knowledge benchmarks
BenchmarkClaude 2.1Gemma 7B
MMLU73.5%66.1%
GPQA Diamond33%—
LMArena Expert—1001
ARC (AI2) Challenge—78.3%
BoolQ—83.2%
OpenBookQA—78.6%
TriviaQA—72.3%

Multilingual Not comparable

Claude 2.1: —, Gemma 7B: 25.1 (#287)

Multilingual benchmarks
BenchmarkClaude 2.1Gemma 7B
LMArena Non-English—999
LMArena Chinese—1035
LMArena French—1025
LMArena Russian—993

Instruction Following Not comparable

Claude 2.1: —, Gemma 7B: 51.5 (#295)

Instruction Following benchmarks
BenchmarkClaude 2.1Gemma 7B
LMArena Instruction Following—1017

Long Context Not comparable

Claude 2.1: —, Gemma 7B: 31.1 (#282)

Long Context benchmarks
BenchmarkClaude 2.1Gemma 7B
LMArena Longer Query—1022

Writing & Preference Not comparable

Claude 2.1: —, Gemma 7B: 27.1 (#302)

Writing & Preference benchmarks
BenchmarkClaude 2.1Gemma 7B
LMArena Text—1056
LMArena Creative Writing—1024
LMArena Multi-Turn—963

Frequently asked questions

Is Claude 2.1 better than Gemma 7B?

Gemma 7B is the stronger model overall, scoring 30.0 to 25.2 on the Noometry Index.

Is Claude 2.1 or Gemma 7B better for coding?

Gemma 7B scores higher on coding benchmarks: 30.5 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Gemma 7B share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Gemma 7B has 27.

Related comparisons

Go deeper