Model comparison

Gemma 7B vs Llama 2-7B

Gemma 7B and Llama 2-7B score almost the same on the Noometry Index (30.0 vs 29.1), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Llama 2-7B Meta

29.1

Rank #317 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemma 7B scores higher in 6 categories and Llama 2-7B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 7B leads 19.9 to 15.7.

Side by side

Gemma 7B and Llama 2-7B specifications
Gemma 7BLlama 2-7B
ProviderGoogleMeta
Noometry Index30.029.1
Released2024-02-212023-07-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2729

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 7B leads

Gemma 7B: 30.5 (#294), Llama 2-7B: 29.2 (#307)

Coding benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Coding10481002
HumanEval+28.7%—
MBPP+43.4%—

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Llama 2-7B: 15.7 (#312)

Reasoning benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Hard Prompts10421009
BIG-Bench Hard55.1%39.2%
Epoch Capabilities Index111.9999.06
HellaSwag82.2%77.2%
PIQA81.2%78.8%
WinoGrande79%69.2%
Chess Puzzles—0%
Adversarial NLI48.7%—
LAMBADA—73.3%

Math Too close to call

Gemma 7B: 31.2 (#228), Llama 2-7B: 30.7 (#233)

Math benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Math10661042
GSM8K46.4%16.7%

Knowledge Too close to call

Gemma 7B: 27.3 (#252), Llama 2-7B: 28.2 (#248)

Knowledge benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Expert10011036
ARC (AI2) Challenge78.3%45.9%
BoolQ83.2%77.9%
MMLU66.1%45.8%
OpenBookQA78.6%58.6%
TriviaQA72.3%73.7%

Multimodal Not comparable

Gemma 7B: —, Llama 2-7B: —

Multimodal benchmarks
BenchmarkGemma 7BLlama 2-7B
ScienceQA—43.1%

Multilingual Gemma 7B leads

Gemma 7B: 25.1 (#287), Llama 2-7B: 23.8 (#293)

Multilingual benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Non-English999973
LMArena Chinese1035973
LMArena French1025970
LMArena Russian993995
LMArena German—978
LMArena Spanish—1007

Instruction Following Too close to call

Gemma 7B: 51.5 (#295), Llama 2-7B: 50.8 (#298)

Instruction Following benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Instruction Following10171006

Long Context Too close to call

Gemma 7B: 31.1 (#282), Llama 2-7B: 30.4 (#287)

Long Context benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Longer Query1022999

Writing & Preference Too close to call

Gemma 7B: 27.1 (#302), Llama 2-7B: 28.0 (#298)

Writing & Preference benchmarks
BenchmarkGemma 7BLlama 2-7B
LMArena Text10561053
LMArena Creative Writing10241033
LMArena Multi-Turn9631029

Frequently asked questions

Is Gemma 7B better than Llama 2-7B?

Gemma 7B and Llama 2-7B score almost the same on the Noometry Index (30.0 vs 29.1), so choose on price, context window or the category you care about most.

Is Gemma 7B or Llama 2-7B better for coding?

Gemma 7B scores higher on coding benchmarks: 30.5 versus 29.2 in the Noometry coding category.

How many benchmarks do Gemma 7B and Llama 2-7B share?

24 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Llama 2-7B has 29.

Related comparisons

Go deeper