Model comparison

Gemma 1.1 7b IT vs Llama2 70b Steerlm Chat

Gemma 1.1 7b IT and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (31.3 vs 31.8), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Gemma 1.1 7b IT Google

31.3

Rank #277 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Gemma 1.1 7b IT scores higher in 4 categories and Llama2 70b Steerlm Chat in 3 categories; 3 gaps are clear of the uncertainty.

Side by side

Gemma 1.1 7b IT and Llama2 70b Steerlm Chat specifications
Gemma 1.1 7b ITLlama2 70b Steerlm Chat
ProviderGoogleNVIDIA
Noometry Index31.331.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked199

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 31.5 (#284), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Coding10841025
HumanEval+35.4%—
MBPP+45%—

Reasoning Too close to call

Gemma 1.1 7b IT: 20.5 (#238), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Hard Prompts10711047

Math Too close to call

Gemma 1.1 7b IT: 32.0 (#220), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Math11071072

Knowledge Not comparable

Gemma 1.1 7b IT: 28.3 (#247), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Expert1039—

Multilingual Too close to call

Gemma 1.1 7b IT: 28.1 (#273), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Non-English10521063
LMArena Chinese1061—
LMArena French1065—
LMArena German1054—
LMArena Japanese971—
LMArena Korean988—
LMArena Russian1046—
LMArena Spanish1049—

Instruction Following Too close to call

Gemma 1.1 7b IT: 54.0 (#283), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Instruction Following10571060

Long Context Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 32.1 (#272), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Longer Query1056998

Writing & Preference Llama2 70b Steerlm Chat leads

Gemma 1.1 7b IT: 30.4 (#288), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkGemma 1.1 7b ITLlama2 70b Steerlm Chat
LMArena Text10941098
LMArena Creative Writing10601091
LMArena Multi-Turn10401058

Frequently asked questions

Is Gemma 1.1 7b IT better than Llama2 70b Steerlm Chat?

Gemma 1.1 7b IT and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (31.3 vs 31.8), so choose on price, context window or the category you care about most.

Is Gemma 1.1 7b IT or Llama2 70b Steerlm Chat better for coding?

Gemma 1.1 7b IT scores higher on coding benchmarks: 31.5 versus 29.9 in the Noometry coding category.

How many benchmarks do Gemma 1.1 7b IT and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Gemma 1.1 7b IT has 19 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper