Model comparison

Gemma 2B vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 29.6 on the Noometry Index.

Last verified . 9 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Gemma 2B scores higher in 0 categories and Llama2 70b Steerlm Chat in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama2 70b Steerlm Chat leads 31.6 to 24.0.

Side by side

Gemma 2B and Llama2 70b Steerlm Chat specifications
Gemma 2BLlama2 70b Steerlm Chat
ProviderGoogleNVIDIA
Noometry Index29.631.8
Released2024-02-21—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked239

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 2B: 29.4 (#305), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Coding10101025
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Llama2 70b Steerlm Chat leads

Gemma 2B: 18.8 (#275), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Hard Prompts9891047
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Llama2 70b Steerlm Chat leads

Gemma 2B: 30.0 (#239), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Math10091072
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multilingual Llama2 70b Steerlm Chat leads

Gemma 2B: 23.0 (#294), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Non-English9581063
LMArena Chinese986—
LMArena Russian937—

Instruction Following Llama2 70b Steerlm Chat leads

Gemma 2B: 48.5 (#302), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Instruction Following9701060

Long Context Too close to call

Gemma 2B: 29.9 (#291), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Longer Query981998

Writing & Preference Llama2 70b Steerlm Chat leads

Gemma 2B: 24.0 (#308), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkGemma 2BLlama2 70b Steerlm Chat
LMArena Text10021098
LMArena Creative Writing9871091
LMArena Multi-Turn9451058

Frequently asked questions

Is Gemma 2B better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 29.6 on the Noometry Index.

Is Gemma 2B or Llama2 70b Steerlm Chat better for coding?

They score almost the same on coding (29.4 vs 29.9); test both on your own repository before choosing.

How many benchmarks do Gemma 2B and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper