Model comparison

Granite 4.2 3b vs Magistral Small

Granite 4.2 3b is the stronger model overall, scoring 39.4 to 30.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 4.2 3b IBM

39.4

Rank #169 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • The widest gap is in reasoning, where Granite 4.2 3b leads 26.0 to 6.8.

Side by side

Granite 4.2 3b and Magistral Small specifications
Granite 4.2 3bMagistral Small
ProviderIBMMistral AI
Noometry Index39.430.2
Released—2025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked1110

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 3b leads

Granite 4.2 3b: 40.0 (#151), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkGranite 4.2 3bMagistral Small
SciCode—35.2%
LMArena Coding1361—

Reasoning Granite 4.2 3b leads

Granite 4.2 3b: 26.0 (#138), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkGranite 4.2 3bMagistral Small
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
LMArena Hard Prompts1306—
DTBench—61.3%
Epoch Capabilities Index—133.19

Math Not comparable

Granite 4.2 3b: —, Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkGranite 4.2 3bMagistral Small
OTIS Mock AIME 2024-2025—30%

Knowledge Granite 4.2 3b leads

Granite 4.2 3b: 36.3 (#171), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkGranite 4.2 3bMagistral Small
GPQA Diamond—56.1%
LMArena Expert1315—

Multilingual Not comparable

Granite 4.2 3b: 42.1 (#198), Magistral Small: —

Multilingual benchmarks
BenchmarkGranite 4.2 3bMagistral Small
LMArena Non-English1268—
LMArena Chinese1269—
LMArena Russian1249—

Instruction Following Not comparable

Granite 4.2 3b: 67.1 (#200), Magistral Small: —

Instruction Following benchmarks
BenchmarkGranite 4.2 3bMagistral Small
LMArena Instruction Following1273—

Long Context Not comparable

Granite 4.2 3b: 39.2 (#185), Magistral Small: —

Long Context benchmarks
BenchmarkGranite 4.2 3bMagistral Small
LMArena Longer Query1291—

Writing & Preference Not comparable

Granite 4.2 3b: 47.2 (#212), Magistral Small: —

Writing & Preference benchmarks
BenchmarkGranite 4.2 3bMagistral Small
LMArena Text1293—
LMArena Creative Writing1205—
LMArena Multi-Turn1290—

Frequently asked questions

Is Granite 4.2 3b better than Magistral Small?

Granite 4.2 3b is the stronger model overall, scoring 39.4 to 30.2 on the Noometry Index.

Is Granite 4.2 3b or Magistral Small better for coding?

Granite 4.2 3b scores higher on coding benchmarks: 40.0 versus 38.4 in the Noometry coding category.

How many benchmarks do Granite 4.2 3b and Magistral Small share?

0 benchmarks have published results for both models. Granite 4.2 3b has 11 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper