Model comparison

Magistral Medium vs Olmo 3.1 32b Instruct

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 35.2 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Magistral Medium scores higher in 0 categories and Olmo 3.1 32b Instruct in 8 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 8.6.

Side by side

Magistral Medium and Olmo 3.1 32b Instruct specifications
Magistral MediumOlmo 3.1 32b Instruct
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index35.239.4
Released2025-03-17—
WeightsOpenOpen
Context window262K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$5—
Results tracked2216

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Magistral Medium: 39.1 (#161), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Coding13191347
SciCode39.2%—

Reasoning Olmo 3.1 32b Instruct leads

Magistral Medium: 8.6 (#348), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Hard Prompts12671322
ARC-AGI-20%—
Kagi LLM Benchmark16.2%—
ARC-AGI-16.1%—
CritPt0.3%—

Math Olmo 3.1 32b Instruct leads

Magistral Medium: 35.1 (#189), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Math12501305

Knowledge Olmo 3.1 32b Instruct leads

Magistral Medium: 33.5 (#202), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Expert12231308

Multilingual Olmo 3.1 32b Instruct leads

Magistral Medium: 39.6 (#224), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Non-English12321275
LMArena Chinese12271304
LMArena French12671328
LMArena German12481282
LMArena Korean11251206
LMArena Russian12241268
LMArena Spanish12711336
LMArena Japanese1175—

Instruction Following Olmo 3.1 32b Instruct leads

Magistral Medium: 66.0 (#211), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Instruction Following12541299

Long Context Too close to call

Magistral Medium: 39.3 (#183), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Longer Query12951312

Writing & Preference Olmo 3.1 32b Instruct leads

Magistral Medium: 46.3 (#219), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkMagistral MediumOlmo 3.1 32b Instruct
LMArena Text12551311
LMArena Creative Writing12451264
LMArena Multi-Turn12751309

Frequently asked questions

Is Magistral Medium better than Olmo 3.1 32b Instruct?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 35.2 on the Noometry Index.

Is Magistral Medium or Olmo 3.1 32b Instruct better for coding?

They score almost the same on coding (39.1 vs 39.5); test both on your own repository before choosing.

How many benchmarks do Magistral Medium and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper