Model comparison

Granite 4.1 8b vs Magistral Medium

Granite 4.1 8b is the stronger model overall, scoring 37.4 to 35.2 on the Noometry Index.

Last verified . 12 shared benchmarks.

Granite 4.1 8b IBM

37.4

Rank #205 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Granite 4.1 8b scores higher in 6 categories and Magistral Medium in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Granite 4.1 8b leads 25.7 to 8.6.

Side by side

Granite 4.1 8b and Magistral Medium specifications
Granite 4.1 8bMagistral Medium
ProviderIBMMistral AI
Noometry Index37.435.2
Released—2025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked1322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Granite 4.1 8b: 30.1 (#297), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Coding13121319
LMArena WebDev1192—
SciCode—39.2%

Reasoning Granite 4.1 8b leads

Granite 4.1 8b: 25.7 (#143), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Hard Prompts12931267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%

Math Granite 4.1 8b leads

Granite 4.1 8b: 36.4 (#166), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Math13121250

Knowledge Granite 4.1 8b leads

Granite 4.1 8b: 36.1 (#174), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Expert13091223

Multilingual Granite 4.1 8b leads

Granite 4.1 8b: 41.7 (#204), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Non-English12611232
LMArena Chinese13371227
LMArena Russian12401224
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Spanish—1271

Instruction Following Too close to call

Granite 4.1 8b: 66.9 (#203), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Instruction Following12691254

Long Context Too close to call

Granite 4.1 8b: 38.7 (#193), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Longer Query12751295

Writing & Preference Granite 4.1 8b leads

Granite 4.1 8b: 48.1 (#204), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkGranite 4.1 8bMagistral Medium
LMArena Text12901255
LMArena Creative Writing12511245
LMArena Multi-Turn12681275

Frequently asked questions

Is Granite 4.1 8b better than Magistral Medium?

Granite 4.1 8b is the stronger model overall, scoring 37.4 to 35.2 on the Noometry Index.

Is Granite 4.1 8b or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 30.1 in the Noometry coding category.

How many benchmarks do Granite 4.1 8b and Magistral Medium share?

12 benchmarks have published results for both models. Granite 4.1 8b has 13 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper