Model comparison

Gemma 1.1 2b IT vs Llama 3.1-70B

Gemma 1.1 2b IT and Llama 3.1-70B score almost the same on the Noometry Index (29.3 vs 29.6), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 2 categories and Llama 3.1-70B in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 1.1 2b IT leads 30.8 to 13.5.

Side by side

Gemma 1.1 2b IT and Llama 3.1-70B specifications
Gemma 1.1 2b ITLlama 3.1-70B
ProviderGoogleMeta
Noometry Index29.329.6
Released—2024-07-23
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.40
Output $ / M tokens—$0.40
Results tracked1635

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 1.1 2b IT: 30.1 (#299), Llama 3.1-70B: 30.3 (#296)

Coding benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Coding10341260
WeirdML—9%
BigCodeBench Instruct—46.1%
BigCodeBench Complete—54.8%
HumanEval+17.7%—
MBPP+23.3%—

Agentic & Tool Use Not comparable

Gemma 1.1 2b IT: —, Llama 3.1-70B: 25.1 (#112)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
TheAgentCompany—6.9%
BALROG—27.9%

Reasoning Llama 3.1-70B leads

Gemma 1.1 2b IT: 19.1 (#270), Llama 3.1-70B: 21.6 (#220)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Hard Prompts10051241
DTBench—60%
LMCA—14.8%
Epoch Capabilities Index—125.92

Math Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.8 (#232), Llama 3.1-70B: 13.5 (#304)

Math benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Math10471252
OTIS Mock AIME 2024-2025—3.6%
Omni-MATH—21%
MATH Level 5—36.7%

Knowledge Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 26.5 (#258), Llama 3.1-70B: 24.2 (#269)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Expert9701209
GPQA Diamond—44.2%
MMLU-Pro—65.3%
GPQA (HELM)—42.6%
MMLU—80.1%

Multilingual Llama 3.1-70B leads

Gemma 1.1 2b IT: 24.6 (#289), Llama 3.1-70B: 38.8 (#225)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Non-English9881219
LMArena Chinese10121215
LMArena German9441222
LMArena Korean8991140
LMArena Russian9901234
LMArena French—1261
LMArena Japanese—1132
LMArena Spanish—1253

Instruction Following Llama 3.1-70B leads

Gemma 1.1 2b IT: 49.9 (#299), Llama 3.1-70B: 65.3 (#223)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Instruction Following9921231
IFEval—82.1%

Long Context Llama 3.1-70B leads

Gemma 1.1 2b IT: 30.6 (#286), Llama 3.1-70B: 37.6 (#214)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Longer Query10031241

Writing & Preference Llama 3.1-70B leads

Gemma 1.1 2b IT: 25.1 (#306), Llama 3.1-70B: 35.4 (#267)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITLlama 3.1-70B
LMArena Text10221261
LMArena Creative Writing9981232
LMArena Multi-Turn9591256
EQ-Bench Creative Writing—784
WildBench—75.8%

Frequently asked questions

Is Gemma 1.1 2b IT better than Llama 3.1-70B?

Gemma 1.1 2b IT and Llama 3.1-70B score almost the same on the Noometry Index (29.3 vs 29.6), so choose on price, context window or the category you care about most.

Is Gemma 1.1 2b IT or Llama 3.1-70B better for coding?

They score almost the same on coding (30.1 vs 30.3); test both on your own repository before choosing.

How many benchmarks do Gemma 1.1 2b IT and Llama 3.1-70B share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Llama 3.1-70B has 35.

Related comparisons

Go deeper