Model comparison

Gemma 1.1 2b IT vs Llama 3-70B

Gemma 1.1 2b IT and Llama 3-70B score almost the same on the Noometry Index (29.3 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 3 categories and Llama 3-70B in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 1.1 2b IT leads 30.8 to 12.8.

Side by side

Gemma 1.1 2b IT and Llama 3-70B specifications
Gemma 1.1 2b ITLlama 3-70B
ProviderGoogleMeta
Noometry Index29.328.8
Released—2024-04-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1631

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

Gemma 1.1 2b IT: 30.1 (#299), Llama 3-70B: 35.8 (#218)

Coding benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Coding10341206
HumanEval+17.7%72%
MBPP+23.3%69%
BigCodeBench Instruct—43.6%
BigCodeBench Complete—54.5%

Agentic & Tool Use Not comparable

Gemma 1.1 2b IT: —, Llama 3-70B: 21.1 (#139)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
Cybench—5%

Reasoning Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 19.1 (#270), Llama 3-70B: 18.0 (#288)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Hard Prompts10051195
Kagi LLM Benchmark—35.1%
DTBench—54.2%
Epoch Capabilities Index—122.93
ForecastBench—57.1
WinoGrande—83.5%

Math Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.8 (#232), Llama 3-70B: 12.8 (#305)

Math benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Math10471218
OTIS Mock AIME 2024-2025—4.3%
MATH Level 5—22.6%

Knowledge Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 26.5 (#258), Llama 3-70B: 20.8 (#277)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Expert9701149
GPQA Diamond—40.6%
MMLU—79.3%

Multilingual Llama 3-70B leads

Gemma 1.1 2b IT: 24.6 (#289), Llama 3-70B: 33.6 (#251)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Non-English9881142
LMArena Chinese10121114
LMArena German9441169
LMArena Korean8991017
LMArena Russian9901159
LMArena French—1232
LMArena Japanese—1017
LMArena Spanish—1241

Instruction Following Llama 3-70B leads

Gemma 1.1 2b IT: 49.9 (#299), Llama 3-70B: 62.5 (#238)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Instruction Following9921194

Long Context Llama 3-70B leads

Gemma 1.1 2b IT: 30.6 (#286), Llama 3-70B: 35.6 (#240)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Longer Query10031174

Writing & Preference Llama 3-70B leads

Gemma 1.1 2b IT: 25.1 (#306), Llama 3-70B: 42.8 (#231)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITLlama 3-70B
LMArena Text10221221
LMArena Creative Writing9981210
LMArena Multi-Turn9591223

Frequently asked questions

Is Gemma 1.1 2b IT better than Llama 3-70B?

Gemma 1.1 2b IT and Llama 3-70B score almost the same on the Noometry Index (29.3 vs 28.8), so choose on price, context window or the category you care about most.

Is Gemma 1.1 2b IT or Llama 3-70B better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Llama 3-70B share?

16 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Llama 3-70B has 31.

Related comparisons

Go deeper