Model comparison

Gemma 1.1 7b IT vs Granite 3.1 2b Instruct

Granite 3.1 2b Instruct is the stronger model overall, scoring 33.2 to 31.3 on the Noometry Index.

Last verified . 12 shared benchmarks.

Gemma 1.1 7b IT Google

31.3

Rank #277 Confirmed

Granite 3.1 2b Instruct IBM

33.2

Rank #247 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Gemma 1.1 7b IT scores higher in 0 categories and Granite 3.1 2b Instruct in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Granite 3.1 2b Instruct leads 57.7 to 54.0.

Side by side

Gemma 1.1 7b IT and Granite 3.1 2b Instruct specifications
Gemma 1.1 7b ITGranite 3.1 2b Instruct
ProviderGoogleIBM
Noometry Index31.333.2
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1912

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 31.5 (#284), Granite 3.1 2b Instruct: 33.4 (#257)

Coding benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Coding10841149
HumanEval+35.4%—
MBPP+45%—

Reasoning Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 20.5 (#238), Granite 3.1 2b Instruct: 22.0 (#209)

Reasoning benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Hard Prompts10711138

Math Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 32.0 (#220), Granite 3.1 2b Instruct: 33.1 (#206)

Math benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Math11071159

Knowledge Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 28.3 (#247), Granite 3.1 2b Instruct: 30.8 (#224)

Knowledge benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Expert10391131

Multilingual Too close to call

Gemma 1.1 7b IT: 28.1 (#273), Granite 3.1 2b Instruct: 29.1 (#269)

Multilingual benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Non-English10521068
LMArena Chinese10611139
LMArena Russian10461063
LMArena French1065—
LMArena German1054—
LMArena Japanese971—
LMArena Korean988—
LMArena Spanish1049—

Instruction Following Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 54.0 (#283), Granite 3.1 2b Instruct: 57.7 (#264)

Instruction Following benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Instruction Following10571116

Long Context Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 32.1 (#272), Granite 3.1 2b Instruct: 35.0 (#244)

Long Context benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Longer Query10561155

Writing & Preference Granite 3.1 2b Instruct leads

Gemma 1.1 7b IT: 30.4 (#288), Granite 3.1 2b Instruct: 34.1 (#274)

Writing & Preference benchmarks
BenchmarkGemma 1.1 7b ITGranite 3.1 2b Instruct
LMArena Text10941127
LMArena Creative Writing10601116
LMArena Multi-Turn10401099

Frequently asked questions

Is Gemma 1.1 7b IT better than Granite 3.1 2b Instruct?

Granite 3.1 2b Instruct is the stronger model overall, scoring 33.2 to 31.3 on the Noometry Index.

Is Gemma 1.1 7b IT or Granite 3.1 2b Instruct better for coding?

Granite 3.1 2b Instruct scores higher on coding benchmarks: 33.4 versus 31.5 in the Noometry coding category.

How many benchmarks do Gemma 1.1 7b IT and Granite 3.1 2b Instruct share?

12 benchmarks have published results for both models. Gemma 1.1 7b IT has 19 scored results on Noometry and Granite 3.1 2b Instruct has 12.

Related comparisons

Go deeper