Model comparison

Gemma 2 27B vs Llama 2-7B

Gemma 2 27B and Llama 2-7B score almost the same on the Noometry Index (29.4 vs 29.1), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Llama 2-7B Meta

29.1

Rank #317 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 2 27B scores higher in 5 categories and Llama 2-7B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 2-7B leads 30.7 to 10.7.

Side by side

Gemma 2 27B and Llama 2-7B specifications
Gemma 2 27BLlama 2-7B
ProviderGoogleMeta
Noometry Index29.429.1
Released2024-06-242023-07-18
WeightsOpenOpen
Context window8K—
Max output2K—
Input $ / M tokens$0.65—
Output $ / M tokens$0.65—
Results tracked3429

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 27B leads

Gemma 2 27B: 34.1 (#246), Llama 2-7B: 29.2 (#307)

Coding benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Coding12111002
BigCodeBench Instruct42.8%—
LiveBench Coding36%—
BigCodeBench Complete52.5%—

Reasoning Too close to call

Gemma 2 27B: 15.3 (#315), Llama 2-7B: 15.7 (#312)

Reasoning benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Hard Prompts11981009
Epoch Capabilities Index122.0899.06
Chess Puzzles—0%
LiveBench Reasoning28.1%—
DTBench48%—
LiveBench Data Analysis47.9%—
LMCA7.1%—
BIG-Bench Hard—39.2%
HellaSwag—77.2%
LAMBADA—73.3%
LiveBench38.2%—
PIQA—78.8%
WinoGrande—69.2%

Math Llama 2-7B leads

Gemma 2 27B: 10.7 (#311), Llama 2-7B: 30.7 (#233)

Math benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Math12121042
OTIS Mock AIME 2024-20251.4%—
LiveBench Math26.5%—
MATH Level 527.9%—
GSM8K—16.7%

Knowledge Llama 2-7B leads

Gemma 2 27B: 19.0 (#280), Llama 2-7B: 28.2 (#248)

Knowledge benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Expert11721036
MMLU75.7%45.8%
GPQA Diamond36.5%—
Confabulations27.1%—
ARC (AI2) Challenge—45.9%
BoolQ—77.9%
OpenBookQA—58.6%
TriviaQA—73.7%

Multimodal Not comparable

Gemma 2 27B: —, Llama 2-7B: —

Multimodal benchmarks
BenchmarkGemma 2 27BLlama 2-7B
ScienceQA—43.1%

Multilingual Gemma 2 27B leads

Gemma 2 27B: 38.6 (#226), Llama 2-7B: 23.8 (#293)

Multilingual benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Non-English1217973
LMArena Chinese1221973
LMArena French1247970
LMArena German1209978
LMArena Russian1234995
LMArena Spanish12281007
LMArena Japanese1175—
LMArena Korean1174—

Instruction Following Gemma 2 27B leads

Gemma 2 27B: 60.5 (#249), Llama 2-7B: 50.8 (#298)

Instruction Following benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Instruction Following12061006
LiveBench Instruction Following58.1%—

Long Context Gemma 2 27B leads

Gemma 2 27B: 37.3 (#218), Llama 2-7B: 30.4 (#287)

Long Context benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Longer Query1231999

Writing & Preference Gemma 2 27B leads

Gemma 2 27B: 44.2 (#225), Llama 2-7B: 28.0 (#298)

Writing & Preference benchmarks
BenchmarkGemma 2 27BLlama 2-7B
LMArena Text12311053
LMArena Creative Writing12411033
LMArena Multi-Turn12241029
LiveBench Language32.6%—

Frequently asked questions

Is Gemma 2 27B better than Llama 2-7B?

Gemma 2 27B and Llama 2-7B score almost the same on the Noometry Index (29.4 vs 29.1), so choose on price, context window or the category you care about most.

Is Gemma 2 27B or Llama 2-7B better for coding?

Gemma 2 27B scores higher on coding benchmarks: 34.1 versus 29.2 in the Noometry coding category.

How many benchmarks do Gemma 2 27B and Llama 2-7B share?

17 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and Llama 2-7B has 29.

Related comparisons

Go deeper