Model comparison

Granite 4.1 8b vs Magistral Small

Granite 4.1 8b is the stronger model overall, scoring 37.4 to 30.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 4.1 8b IBM

37.4

Rank #205 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • The widest gap is in reasoning, where Granite 4.1 8b leads 25.7 to 6.8.

Side by side

Granite 4.1 8b and Magistral Small specifications
Granite 4.1 8bMagistral Small
ProviderIBMMistral AI
Noometry Index37.430.2
Released—2025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked1310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Granite 4.1 8b: 30.1 (#297), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkGranite 4.1 8bMagistral Small
LMArena WebDev1192—
SciCode—35.2%
LMArena Coding1312—

Reasoning Granite 4.1 8b leads

Granite 4.1 8b: 25.7 (#143), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkGranite 4.1 8bMagistral Small
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
LMArena Hard Prompts1293—
DTBench—61.3%
Epoch Capabilities Index—133.19

Math Granite 4.1 8b leads

Granite 4.1 8b: 36.4 (#166), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkGranite 4.1 8bMagistral Small
OTIS Mock AIME 2024-2025—30%
LMArena Math1312—

Knowledge Granite 4.1 8b leads

Granite 4.1 8b: 36.1 (#174), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkGranite 4.1 8bMagistral Small
GPQA Diamond—56.1%
LMArena Expert1309—

Multilingual Not comparable

Granite 4.1 8b: 41.7 (#204), Magistral Small: —

Multilingual benchmarks
BenchmarkGranite 4.1 8bMagistral Small
LMArena Non-English1261—
LMArena Chinese1337—
LMArena Russian1240—

Instruction Following Not comparable

Granite 4.1 8b: 66.9 (#203), Magistral Small: —

Instruction Following benchmarks
BenchmarkGranite 4.1 8bMagistral Small
LMArena Instruction Following1269—

Long Context Not comparable

Granite 4.1 8b: 38.7 (#193), Magistral Small: —

Long Context benchmarks
BenchmarkGranite 4.1 8bMagistral Small
LMArena Longer Query1275—

Writing & Preference Not comparable

Granite 4.1 8b: 48.1 (#204), Magistral Small: —

Writing & Preference benchmarks
BenchmarkGranite 4.1 8bMagistral Small
LMArena Text1290—
LMArena Creative Writing1251—
LMArena Multi-Turn1268—

Frequently asked questions

Is Granite 4.1 8b better than Magistral Small?

Granite 4.1 8b is the stronger model overall, scoring 37.4 to 30.2 on the Noometry Index.

Is Granite 4.1 8b or Magistral Small better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 30.1 in the Noometry coding category.

How many benchmarks do Granite 4.1 8b and Magistral Small share?

0 benchmarks have published results for both models. Granite 4.1 8b has 13 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper