Model comparison

Granite 3.0 2b Instruct vs Mistral Nemo

Granite 3.0 2b Instruct is the stronger model overall, scoring 30.8 to 26.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 3.0 2b Instruct IBM

30.8

Rank #286 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • The widest gap is in knowledge, where Granite 3.0 2b Instruct leads 29.0 to 12.3.

Side by side

Granite 3.0 2b Instruct and Mistral Nemo specifications
Granite 3.0 2b InstructMistral Nemo
ProviderIBMMistral AI
Noometry Index30.826.4
Released—2024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked1310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 3.0 2b Instruct: 28.3 (#316), Mistral Nemo: —

Coding benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
BigCodeBench Instruct20.5%—
LMArena Coding1090—

Agentic & Tool Use Not comparable

Granite 3.0 2b Instruct: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Too close to call

Granite 3.0 2b Instruct: 20.5 (#235), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
LMArena Hard Prompts1073—
DTBench—48.6%
Epoch Capabilities Index—118.68
PIQA—83.5%

Math Granite 3.0 2b Instruct leads

Granite 3.0 2b Instruct: 32.2 (#217), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
LMArena Math1117—
MATH Level 5—10.8%
GSM8K—84.2%

Knowledge Granite 3.0 2b Instruct leads

Granite 3.0 2b Instruct: 29.0 (#241), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
GPQA Diamond—29.9%
LMArena Expert1064—
BoolQ—82.5%

Multilingual Not comparable

Granite 3.0 2b Instruct: 27.0 (#278), Mistral Nemo: —

Multilingual benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
LMArena Non-English1033—
LMArena Chinese1070—
LMArena Russian1045—

Instruction Following Not comparable

Granite 3.0 2b Instruct: 53.9 (#284), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
LMArena Instruction Following1056—

Long Context Not comparable

Granite 3.0 2b Instruct: 32.5 (#268), Mistral Nemo: —

Long Context benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
LMArena Longer Query1070—

Writing & Preference Granite 3.0 2b Instruct leads

Granite 3.0 2b Instruct: 29.6 (#292), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkGranite 3.0 2b InstructMistral Nemo
LMArena Text1080—
LMArena Creative Writing1046—
EQ-Bench Creative Writing—881
LMArena Multi-Turn1053—

Frequently asked questions

Is Granite 3.0 2b Instruct better than Mistral Nemo?

Granite 3.0 2b Instruct is the stronger model overall, scoring 30.8 to 26.4 on the Noometry Index.

How many benchmarks do Granite 3.0 2b Instruct and Mistral Nemo share?

0 benchmarks have published results for both models. Granite 3.0 2b Instruct has 13 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper