Model comparison

Gemma 2 27B vs Llama 3.2 3B

Gemma 2 27B and Llama 3.2 3B score almost the same on the Noometry Index (29.4 vs 28.9), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Gemma 2 27B scores higher in 5 categories and Llama 3.2 3B in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 3.2 3B leads 32.4 to 10.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 52.5% for Gemma 2 27B and 28.3% for Llama 3.2 3B.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.65 / $0.65 for Gemma 2 27B.
  • Llama 3.2 3B accepts more context: 131K tokens versus 8K.

Side by side

Gemma 2 27B and Llama 3.2 3B specifications
Gemma 2 27BLlama 3.2 3B
ProviderGoogleMeta
Noometry Index29.428.9
Released2024-06-242024-09-24
WeightsOpenOpen
Context window8K131K
Max output2K118K
Input $ / M tokens$0.65$0.05
Output $ / M tokens$0.65$0.33
Results tracked3418

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 27B leads

Gemma 2 27B: 34.1 (#246), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
BigCodeBench Instruct42.8%23.4%
LMArena Coding12111098
BigCodeBench Complete52.5%28.3%
LiveBench Coding36%—

Agentic & Tool Use Not comparable

Gemma 2 27B: —, Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
Berkeley Function Calling Leaderboard—21.9%
BALROG—10.1%

Reasoning Llama 3.2 3B leads

Gemma 2 27B: 15.3 (#315), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Hard Prompts11981095
LiveBench Reasoning28.1%—
DTBench48%—
LiveBench Data Analysis47.9%—
LMCA7.1%—
Epoch Capabilities Index122.08—
LiveBench38.2%—

Math Llama 3.2 3B leads

Gemma 2 27B: 10.7 (#311), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Math12121126
OTIS Mock AIME 2024-20251.4%—
LiveBench Math26.5%—
MATH Level 527.9%—

Knowledge Llama 3.2 3B leads

Gemma 2 27B: 19.0 (#280), Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Expert11721090
GPQA Diamond36.5%—
Confabulations27.1%—
MMLU75.7%—

Multilingual Gemma 2 27B leads

Gemma 2 27B: 38.6 (#226), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Non-English12171019
LMArena Chinese12211017
LMArena German12091056
LMArena Russian1234949
LMArena French1247—
LMArena Japanese1175—
LMArena Korean1174—
LMArena Spanish1228—

Instruction Following Gemma 2 27B leads

Gemma 2 27B: 60.5 (#249), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Instruction Following12061089
LiveBench Instruction Following58.1%—

Long Context Gemma 2 27B leads

Gemma 2 27B: 37.3 (#218), Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Longer Query12311100

Writing & Preference Gemma 2 27B leads

Gemma 2 27B: 44.2 (#225), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkGemma 2 27BLlama 3.2 3B
LMArena Text12311110
LMArena Creative Writing12411094
LMArena Multi-Turn12241105
EQ-Bench Creative Writing—595
LiveBench Language32.6%—

Frequently asked questions

Is Gemma 2 27B better than Llama 3.2 3B?

Gemma 2 27B and Llama 3.2 3B score almost the same on the Noometry Index (29.4 vs 28.9), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 2 27B or Llama 3.2 3B?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Gemma 2 27B lists at $0.65 and $0.65.

Is Gemma 2 27B or Llama 3.2 3B better for coding?

Gemma 2 27B scores higher on coding benchmarks: 34.1 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 8K.

How many benchmarks do Gemma 2 27B and Llama 3.2 3B share?

15 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper