Model comparison

Gemma 2 2b IT vs Llama 3.2 1B

Gemma 2 2b IT is the stronger model overall, scoring 33.1 to 20.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemma 2 2b IT Google

33.1

Rank #248 Confirmed

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemma 2 2b IT scores higher in 8 categories and Llama 3.2 1B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemma 2 2b IT leads 29.9 to 7.2.

Side by side

Gemma 2 2b IT and Llama 3.2 1B specifications
Gemma 2 2b ITLlama 3.2 1B
ProviderGoogleMeta
Noometry Index33.120.1
Released—2024-09-24
WeightsOpenOpen
Context window—60K
Max output—54K
Input $ / M tokens—$0.027
Output $ / M tokens—$0.20
Results tracked1722

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 2b IT leads

Gemma 2 2b IT: 32.3 (#273), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Coding11121070
BigCodeBench Instruct—8.2%
BigCodeBench Complete—11.3%

Agentic & Tool Use Not comparable

Gemma 2 2b IT: —, Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
Berkeley Function Calling Leaderboard—10.8%
BALROG—6.6%

Reasoning Gemma 2 2b IT leads

Gemma 2 2b IT: 21.4 (#224), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Hard Prompts11131044
Chess Puzzles—0%
Epoch Capabilities Index—101.99

Math Gemma 2 2b IT leads

Gemma 2 2b IT: 32.6 (#212), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Math11351086
OTIS Mock AIME 2024-2025—0.6%

Knowledge Gemma 2 2b IT leads

Gemma 2 2b IT: 29.9 (#231), Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Expert10961007
GPQA Diamond—23.9%

Multilingual Gemma 2 2b IT leads

Gemma 2 2b IT: 32.3 (#257), Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Non-English1121973
LMArena Chinese1132959
LMArena German11141014
LMArena Russian1118941
LMArena French1157—
LMArena Japanese1083—
LMArena Korean1055—
LMArena Spanish1140—

Instruction Following Gemma 2 2b IT leads

Gemma 2 2b IT: 57.8 (#263), Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Instruction Following11181031

Long Context Gemma 2 2b IT leads

Gemma 2 2b IT: 34.3 (#250), Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Longer Query11301050

Writing & Preference Gemma 2 2b IT leads

Gemma 2 2b IT: 36.5 (#263), Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkGemma 2 2b ITLlama 3.2 1B
LMArena Text11561055
LMArena Creative Writing11471033
LMArena Multi-Turn11181030
EQ-Bench Creative Writing—200

Frequently asked questions

Is Gemma 2 2b IT better than Llama 3.2 1B?

Gemma 2 2b IT is the stronger model overall, scoring 33.1 to 20.1 on the Noometry Index.

Is Gemma 2 2b IT or Llama 3.2 1B better for coding?

Gemma 2 2b IT scores higher on coding benchmarks: 32.3 versus 21.1 in the Noometry coding category.

How many benchmarks do Gemma 2 2b IT and Llama 3.2 1B share?

13 benchmarks have published results for both models. Gemma 2 2b IT has 17 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper