Model comparison

Gemma 7B vs Llama 3.1-405B

Gemma 7B and Llama 3.1-405B score almost the same on the Noometry Index (30.0 vs 30.7), so choose on price, context window or the category you care about most.

Last verified . 21 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemma 7B scores higher in 2 categories and Llama 3.1-405B in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Llama 3.1-405B leads 40.7 to 25.1.

Side by side

Gemma 7B and Llama 3.1-405B specifications
Gemma 7BLlama 3.1-405B
ProviderGoogleMeta
Noometry Index30.030.7
Released2024-02-212024-07-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2742

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Gemma 7B: 30.5 (#294), Llama 3.1-405B: 33.1 (#262)

Coding benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Coding10481291
WeirdML—21.4%
HumanEval+28.7%—
MBPP+43.4%—

Agentic & Tool Use Not comparable

Gemma 7B: —, Llama 3.1-405B: 21.0 (#140)

Agentic & Tool Use benchmarks
BenchmarkGemma 7BLlama 3.1-405B
TheAgentCompany—7.4%
Cybench—7.5%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Llama 3.1-405B: 16.8 (#300)

Reasoning benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Hard Prompts10421269
BIG-Bench Hard55.1%82.9%
Epoch Capabilities Index111.99128.75
HellaSwag82.2%89.2%
PIQA81.2%85.9%
WinoGrande79%89.2%
SimpleBench—23%
Kagi LLM Benchmark—45%
DTBench—61.4%
Adversarial NLI48.7%—
ForecastBench—59.9

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Llama 3.1-405B: 18.4 (#290)

Math benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Math10661281
OTIS Mock AIME 2024-2025—9.7%
Omni-MATH—24.9%
MATH Level 5—49.8%
GSM8K46.4%—

Knowledge Llama 3.1-405B leads

Gemma 7B: 27.3 (#252), Llama 3.1-405B: 30.4 (#227)

Knowledge benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Expert10011243
ARC (AI2) Challenge78.3%95.3%
MMLU66.1%84.5%
TriviaQA72.3%82.7%
GPQA Diamond—50.9%
MMLU-Pro—72.3%
Confabulations—17.6%
GPQA (HELM)—52.2%
BoolQ83.2%—
OpenBookQA78.6%—

Multilingual Llama 3.1-405B leads

Gemma 7B: 25.1 (#287), Llama 3.1-405B: 40.7 (#214)

Multilingual benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Non-English9991248
LMArena Chinese10351242
LMArena French10251279
LMArena Russian9931265
LMArena German—1252
LMArena Japanese—1208
LMArena Korean—1184
LMArena Spanish—1260

Instruction Following Llama 3.1-405B leads

Gemma 7B: 51.5 (#295), Llama 3.1-405B: 65.9 (#214)

Instruction Following benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Instruction Following10171259
IFEval—81.1%

Long Context Llama 3.1-405B leads

Gemma 7B: 31.1 (#282), Llama 3.1-405B: 38.4 (#197)

Long Context benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Longer Query10221266

Writing & Preference Llama 3.1-405B leads

Gemma 7B: 27.1 (#302), Llama 3.1-405B: 38.9 (#251)

Writing & Preference benchmarks
BenchmarkGemma 7BLlama 3.1-405B
LMArena Text10561284
LMArena Creative Writing10241262
LMArena Multi-Turn9631297
EQ-Bench Creative Writing—870
WildBench—78.3%

Frequently asked questions

Is Gemma 7B better than Llama 3.1-405B?

Gemma 7B and Llama 3.1-405B score almost the same on the Noometry Index (30.0 vs 30.7), so choose on price, context window or the category you care about most.

Is Gemma 7B or Llama 3.1-405B better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 30.5 in the Noometry coding category.

How many benchmarks do Gemma 7B and Llama 3.1-405B share?

21 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Llama 3.1-405B has 42.

Related comparisons

Go deeper