Model comparison

Codellama 70b Instruct vs Gemma 1.1 2b IT

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 29.3 on the Noometry Index.

Last verified . 5 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Codellama 70b Instruct scores higher in 5 categories and Gemma 1.1 2b IT in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Codellama 70b Instruct leads 33.4 to 25.1.

Side by side

Codellama 70b Instruct and Gemma 1.1 2b IT specifications
Codellama 70b InstructGemma 1.1 2b IT
ProviderMetaGoogle
Noometry Index33.729.3
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked716

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Gemma 1.1 2b IT: 30.1 (#299)

Coding benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
HumanEval+65.9%17.7%
BigCodeBench Instruct40.7%—
LMArena Coding—1034
BigCodeBench Complete49.6%—
MBPP+—23.3%

Reasoning Too close to call

Codellama 70b Instruct: 20.1 (#242), Gemma 1.1 2b IT: 19.1 (#270)

Reasoning benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Hard Prompts10521005

Math Not comparable

Codellama 70b Instruct: —, Gemma 1.1 2b IT: 30.8 (#232)

Math benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Math—1047

Knowledge Not comparable

Codellama 70b Instruct: —, Gemma 1.1 2b IT: 26.5 (#258)

Knowledge benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Expert—970

Multilingual Too close to call

Codellama 70b Instruct: 24.8 (#288), Gemma 1.1 2b IT: 24.6 (#289)

Multilingual benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Non-English992988
LMArena Chinese—1012
LMArena German—944
LMArena Korean—899
LMArena Russian—990

Instruction Following Codellama 70b Instruct leads

Codellama 70b Instruct: 51.9 (#293), Gemma 1.1 2b IT: 49.9 (#299)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Instruction Following1024992

Long Context Not comparable

Codellama 70b Instruct: —, Gemma 1.1 2b IT: 30.6 (#286)

Long Context benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Longer Query—1003

Writing & Preference Codellama 70b Instruct leads

Codellama 70b Instruct: 33.4 (#277), Gemma 1.1 2b IT: 25.1 (#306)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructGemma 1.1 2b IT
LMArena Text10571022
LMArena Creative Writing—998
LMArena Multi-Turn—959

Frequently asked questions

Is Codellama 70b Instruct better than Gemma 1.1 2b IT?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 29.3 on the Noometry Index.

Is Codellama 70b Instruct or Gemma 1.1 2b IT better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 30.1 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Gemma 1.1 2b IT share?

5 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Gemma 1.1 2b IT has 16.

Related comparisons

Go deeper