Model comparison

Gemma 7B vs Llama 3-8B

Gemma 7B is the stronger model overall, scoring 30.0 to 25.5 on the Noometry Index.

Last verified . 22 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Gemma 7B scores higher in 3 categories and Llama 3-8B in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 7B leads 31.2 to 8.8.

Side by side

Gemma 7B and Llama 3-8B specifications
Gemma 7BLlama 3-8B
ProviderGoogleMeta
Noometry Index30.025.5
Released2024-02-212024-04-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2734

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 7B: 30.5 (#294), Llama 3-8B: 31.0 (#289)

Coding benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Coding10481152
HumanEval+28.7%56.7%
MBPP+43.4%54.8%
BigCodeBench Instruct—31.9%
BigCodeBench Complete—36.9%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Llama 3-8B: 14.3 (#326)

Reasoning benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Hard Prompts10421133
Adversarial NLI48.7%57.3%
Epoch Capabilities Index111.99116.45
WinoGrande79%75.7%
Chess Puzzles—0%
DTBench—43.9%
BIG-Bench Hard55.1%—
ForecastBench—58.6
HellaSwag82.2%—
PIQA81.2%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Llama 3-8B: 8.8 (#323)

Math benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Math10661151
OTIS Mock AIME 2024-2025—1.9%
MATH Level 5—6.1%
GSM8K46.4%—

Knowledge Gemma 7B leads

Gemma 7B: 27.3 (#252), Llama 3-8B: 7.8 (#308)

Knowledge benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Expert10011113
ARC (AI2) Challenge78.3%82.8%
MMLU66.1%68.8%
OpenBookQA78.6%82.6%
TriviaQA72.3%67.7%
GPQA Diamond—26.1%
BoolQ83.2%—

Multilingual Llama 3-8B leads

Gemma 7B: 25.1 (#287), Llama 3-8B: 30.8 (#261)

Multilingual benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Non-English9991098
LMArena Chinese10351076
LMArena French10251159
LMArena Russian9931109
LMArena German—1104
LMArena Japanese—967
LMArena Korean—1004
LMArena Spanish—1173

Instruction Following Llama 3-8B leads

Gemma 7B: 51.5 (#295), Llama 3-8B: 58.4 (#260)

Instruction Following benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Instruction Following10171127

Long Context Llama 3-8B leads

Gemma 7B: 31.1 (#282), Llama 3-8B: 34.2 (#251)

Long Context benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Longer Query10221128

Writing & Preference Llama 3-8B leads

Gemma 7B: 27.1 (#302), Llama 3-8B: 37.5 (#256)

Writing & Preference benchmarks
BenchmarkGemma 7BLlama 3-8B
LMArena Text10561166
LMArena Creative Writing10241150
LMArena Multi-Turn9631152

Frequently asked questions

Is Gemma 7B better than Llama 3-8B?

Gemma 7B is the stronger model overall, scoring 30.0 to 25.5 on the Noometry Index.

Is Gemma 7B or Llama 3-8B better for coding?

They score almost the same on coding (30.5 vs 31.0); test both on your own repository before choosing.

How many benchmarks do Gemma 7B and Llama 3-8B share?

22 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Llama 3-8B has 34.

Related comparisons

Go deeper