Model comparison

Gemma 7B vs Llama 3.1-8B

Gemma 7B is the stronger model overall, scoring 30.0 to 23.0 on the Noometry Index.

Last verified . 20 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Llama 3.1-8B Meta

23.0

Rank #352 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemma 7B scores higher in 4 categories and Llama 3.1-8B in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 7B leads 31.2 to 10.2.

Side by side

Gemma 7B and Llama 3.1-8B specifications
Gemma 7BLlama 3.1-8B
ProviderGoogleMeta
Noometry Index30.023.0
Released2024-02-212024-07-23
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.08
Results tracked2743

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 7B leads

Gemma 7B: 30.5 (#294), Llama 3.1-8B: 20.2 (#340)

Coding benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Coding10481195
HumanEval+28.7%62.8%
MBPP+43.4%55.6%
SciCode—13.2%
WeirdML—1.7%
BigCodeBench Instruct—32.8%
BigCodeBench Complete—40.5%

Agentic & Tool Use Not comparable

Gemma 7B: —, Llama 3.1-8B: 22.5 (#131)

Agentic & Tool Use benchmarks
BenchmarkGemma 7BLlama 3.1-8B
Berkeley Function Calling Leaderboard—25.8%
BALROG—15.1%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Llama 3.1-8B: 14.9 (#321)

Reasoning benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Hard Prompts10421175
Epoch Capabilities Index111.99116.57
PIQA81.2%81.2%
CritPt—0%
Chess Puzzles—0%
DTBench—50.9%
LMCA—5.4%
Adversarial NLI48.7%—
BIG-Bench Hard55.1%—
HellaSwag82.2%—
WinoGrande79%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Llama 3.1-8B: 10.2 (#317)

Math benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Math10661179
GSM8K46.4%82.4%
OTIS Mock AIME 2024-2025—1.7%
Omni-MATH—13.7%
MATH Level 5—22.9%

Knowledge Gemma 7B leads

Gemma 7B: 27.3 (#252), Llama 3.1-8B: 8.0 (#307)

Knowledge benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Expert10011144
BoolQ83.2%82.8%
MMLU66.1%56.1%
GPQA Diamond—27%
MMLU-Pro—40.6%
GPQA (HELM)—24.7%
ARC (AI2) Challenge78.3%—
OpenBookQA78.6%—
TriviaQA72.3%—

Multilingual Llama 3.1-8B leads

Gemma 7B: 25.1 (#287), Llama 3.1-8B: 34.0 (#249)

Multilingual benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Non-English9991148
LMArena Chinese10351151
LMArena French10251177
LMArena Russian9931158
LMArena German—1144
LMArena Japanese—1061
LMArena Korean—1053
LMArena Spanish—1169

Instruction Following Llama 3.1-8B leads

Gemma 7B: 51.5 (#295), Llama 3.1-8B: 58.9 (#258)

Instruction Following benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Instruction Following10171159
IFEval—74.3%

Long Context Llama 3.1-8B leads

Gemma 7B: 31.1 (#282), Llama 3.1-8B: 35.8 (#238)

Long Context benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Longer Query10221182

Writing & Preference Llama 3.1-8B leads

Gemma 7B: 27.1 (#302), Llama 3.1-8B: 29.7 (#290)

Writing & Preference benchmarks
BenchmarkGemma 7BLlama 3.1-8B
LMArena Text10561187
LMArena Creative Writing10241154
LMArena Multi-Turn9631172
EQ-Bench Creative Writing—713
WildBench—68.7%

Frequently asked questions

Is Gemma 7B better than Llama 3.1-8B?

Gemma 7B is the stronger model overall, scoring 30.0 to 23.0 on the Noometry Index.

Is Gemma 7B or Llama 3.1-8B better for coding?

Gemma 7B scores higher on coding benchmarks: 30.5 versus 20.2 in the Noometry coding category.

How many benchmarks do Gemma 7B and Llama 3.1-8B share?

20 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Llama 3.1-8B has 43.

Related comparisons

Go deeper