Model comparison

Magistral Medium vs Olmo 3.1 32b Think

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 35.2 on the Noometry Index.

Last verified . 15 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Magistral Medium scores higher in 5 categories and Olmo 3.1 32b Think in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Think leads 25.2 to 8.6.

Side by side

Magistral Medium and Olmo 3.1 32b Think specifications
Magistral MediumOlmo 3.1 32b Think
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index35.237.9
Released2025-03-17—
WeightsOpenOpen
Context window262K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$5—
Results tracked2215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Magistral Medium: 39.1 (#161), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Coding13191291
SciCode39.2%—

Reasoning Olmo 3.1 32b Think leads

Magistral Medium: 8.6 (#348), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Hard Prompts12671272
ARC-AGI-20%—
Kagi LLM Benchmark16.2%—
ARC-AGI-16.1%—
CritPt0.3%—

Math Olmo 3.1 32b Think leads

Magistral Medium: 35.1 (#189), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Math12501305

Knowledge Olmo 3.1 32b Think leads

Magistral Medium: 33.5 (#202), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Expert12231295

Multilingual Magistral Medium leads

Magistral Medium: 39.6 (#224), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Non-English12321209
LMArena Chinese12271242
LMArena French12671260
LMArena German12481262
LMArena Russian12241193
LMArena Spanish12711289
LMArena Japanese1175—
LMArena Korean1125—

Instruction Following Too close to call

Magistral Medium: 66.0 (#211), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Instruction Following12541247

Long Context Too close to call

Magistral Medium: 39.3 (#183), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Longer Query12951272

Writing & Preference Too close to call

Magistral Medium: 46.3 (#219), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Think
LMArena Text12551272
LMArena Creative Writing12451226
LMArena Multi-Turn12751252

Frequently asked questions

Is Magistral Medium better than Olmo 3.1 32b Think?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 35.2 on the Noometry Index.

Is Magistral Medium or Olmo 3.1 32b Think better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 37.7 in the Noometry coding category.

How many benchmarks do Magistral Medium and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper