Model comparison

Gemma 2B vs Llama 3.2 3B

Gemma 2B and Llama 3.2 3B score almost the same on the Noometry Index (29.6 vs 28.9), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Gemma 2B scores higher in 1 category and Llama 3.2 3B in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.2 3B leads 56.0 to 48.5.

Side by side

Gemma 2B and Llama 3.2 3B specifications
Gemma 2BLlama 3.2 3B
ProviderGoogleMeta
Noometry Index29.628.9
Released2024-02-212024-09-24
WeightsOpenOpen
Context window—131K
Max output—118K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.33
Results tracked2318

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2B leads

Gemma 2B: 29.4 (#305), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Coding10101098
BigCodeBench Instruct—23.4%
BigCodeBench Complete—28.3%
HumanEval+20.7%—
MBPP+34.1%—

Agentic & Tool Use Not comparable

Gemma 2B: —, Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkGemma 2BLlama 3.2 3B
Berkeley Function Calling Leaderboard—21.9%
BALROG—10.1%

Reasoning Llama 3.2 3B leads

Gemma 2B: 18.8 (#275), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Hard Prompts9891095
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Llama 3.2 3B leads

Gemma 2B: 30.0 (#239), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Math10091126
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Expert—1090
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multilingual Llama 3.2 3B leads

Gemma 2B: 23.0 (#294), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Non-English9581019
LMArena Chinese9861017
LMArena Russian937949
LMArena German—1056

Instruction Following Llama 3.2 3B leads

Gemma 2B: 48.5 (#302), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Instruction Following9701089

Long Context Llama 3.2 3B leads

Gemma 2B: 29.9 (#291), Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Longer Query9811100

Writing & Preference Too close to call

Gemma 2B: 24.0 (#308), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkGemma 2BLlama 3.2 3B
LMArena Text10021110
LMArena Creative Writing9871094
LMArena Multi-Turn9451105
EQ-Bench Creative Writing—595

Frequently asked questions

Is Gemma 2B better than Llama 3.2 3B?

Gemma 2B and Llama 3.2 3B score almost the same on the Noometry Index (29.6 vs 28.9), so choose on price, context window or the category you care about most.

Is Gemma 2B or Llama 3.2 3B better for coding?

Gemma 2B scores higher on coding benchmarks: 29.4 versus 27.6 in the Noometry coding category.

How many benchmarks do Gemma 2B and Llama 3.2 3B share?

11 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper