Model comparison

Granite 3.1 2b Instruct vs Granite 4.0 Micro

Granite 3.1 2b Instruct is the stronger model overall, scoring 33.2 to 29.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 3.1 2b Instruct IBM

33.2

Rank #247 Confirmed

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Summary

  • The widest gap is in math, where Granite 3.1 2b Instruct leads 33.1 to 12.0.

Side by side

Granite 3.1 2b Instruct and Granite 4.0 Micro specifications
Granite 3.1 2b InstructGranite 4.0 Micro
ProviderIBMIBM
Noometry Index33.229.0
Released—2025-10-02
WeightsOpenOpen
Context window—131K
Max output—118K
Input $ / M tokens—$0.017
Output $ / M tokens—$0.11
Results tracked128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 3.1 2b Instruct: 33.4 (#257), Granite 4.0 Micro: —

Coding benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
LMArena Coding1149—

Reasoning Granite 3.1 2b Instruct leads

Granite 3.1 2b Instruct: 22.0 (#209), Granite 4.0 Micro: 19.2 (#265)

Reasoning benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
Chess Puzzles—0%
LMArena Hard Prompts1138—

Math Granite 3.1 2b Instruct leads

Granite 3.1 2b Instruct: 33.1 (#206), Granite 4.0 Micro: 12.0 (#307)

Math benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
OTIS Mock AIME 2024-2025—2.8%
Omni-MATH—20.9%
LMArena Math1159—

Knowledge Granite 3.1 2b Instruct leads

Granite 3.1 2b Instruct: 30.8 (#224), Granite 4.0 Micro: 9.9 (#304)

Knowledge benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
GPQA Diamond—28.3%
MMLU-Pro—39.5%
GPQA (HELM)—30.7%
LMArena Expert1131—

Multilingual Not comparable

Granite 3.1 2b Instruct: 29.1 (#269), Granite 4.0 Micro: —

Multilingual benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
LMArena Non-English1068—
LMArena Chinese1139—
LMArena Russian1063—

Instruction Following Granite 4.0 Micro leads

Granite 3.1 2b Instruct: 57.7 (#264), Granite 4.0 Micro: 69.9 (#169)

Instruction Following benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
IFEval—84.9%
LMArena Instruction Following1116—

Long Context Not comparable

Granite 3.1 2b Instruct: 35.0 (#244), Granite 4.0 Micro: —

Long Context benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
LMArena Longer Query1155—

Writing & Preference Granite 4.0 Micro leads

Granite 3.1 2b Instruct: 34.1 (#274), Granite 4.0 Micro: 46.7 (#216)

Writing & Preference benchmarks
BenchmarkGranite 3.1 2b InstructGranite 4.0 Micro
LMArena Text1127—
LMArena Creative Writing1116—
WildBench—67%
LMArena Multi-Turn1099—

Frequently asked questions

Is Granite 3.1 2b Instruct better than Granite 4.0 Micro?

Granite 3.1 2b Instruct is the stronger model overall, scoring 33.2 to 29.0 on the Noometry Index.

How many benchmarks do Granite 3.1 2b Instruct and Granite 4.0 Micro share?

0 benchmarks have published results for both models. Granite 3.1 2b Instruct has 12 scored results on Noometry and Granite 4.0 Micro has 8.

Related comparisons

Go deeper