Model comparison

Gemma 1.1 7b IT vs Llama 3.2 3B

Gemma 1.1 7b IT is the stronger model overall, scoring 31.3 to 28.9 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemma 1.1 7b IT Google

31.3

Rank #277 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemma 1.1 7b IT scores higher in 3 categories and Llama 3.2 3B in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemma 1.1 7b IT leads 30.4 to 24.7.

Side by side

Gemma 1.1 7b IT and Llama 3.2 3B specifications
Gemma 1.1 7b ITLlama 3.2 3B
ProviderGoogleMeta
Noometry Index31.328.9
Released—2024-09-24
WeightsOpenOpen
Context window—131K
Max output—118K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.33
Results tracked1918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 31.5 (#284), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Coding10841098
BigCodeBench Instruct—23.4%
BigCodeBench Complete—28.3%
HumanEval+35.4%—
MBPP+45%—

Agentic & Tool Use Not comparable

Gemma 1.1 7b IT: —, Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
Berkeley Function Calling Leaderboard—21.9%
BALROG—10.1%

Reasoning Too close to call

Gemma 1.1 7b IT: 20.5 (#238), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Hard Prompts10711095

Math Too close to call

Gemma 1.1 7b IT: 32.0 (#220), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Math11071126

Knowledge Llama 3.2 3B leads

Gemma 1.1 7b IT: 28.3 (#247), Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Expert10391090

Multilingual Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 28.1 (#273), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Non-English10521019
LMArena Chinese10611017
LMArena German10541056
LMArena Russian1046949
LMArena French1065—
LMArena Japanese971—
LMArena Korean988—
LMArena Spanish1049—

Instruction Following Llama 3.2 3B leads

Gemma 1.1 7b IT: 54.0 (#283), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Instruction Following10571089

Long Context Llama 3.2 3B leads

Gemma 1.1 7b IT: 32.1 (#272), Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Longer Query10561100

Writing & Preference Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 30.4 (#288), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkGemma 1.1 7b ITLlama 3.2 3B
LMArena Text10941110
LMArena Creative Writing10601094
LMArena Multi-Turn10401105
EQ-Bench Creative Writing—595

Frequently asked questions

Is Gemma 1.1 7b IT better than Llama 3.2 3B?

Gemma 1.1 7b IT is the stronger model overall, scoring 31.3 to 28.9 on the Noometry Index.

Is Gemma 1.1 7b IT or Llama 3.2 3B better for coding?

Gemma 1.1 7b IT scores higher on coding benchmarks: 31.5 versus 27.6 in the Noometry coding category.

How many benchmarks do Gemma 1.1 7b IT and Llama 3.2 3B share?

13 benchmarks have published results for both models. Gemma 1.1 7b IT has 19 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper