Model comparison

Magistral Medium vs Mistral Small 3.2

Magistral Medium is the stronger model overall, scoring 35.2 to 31.2 on the Noometry Index. Mistral Small 3.2 costs 21× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Mistral Small 3.2 Mistral AI

31.2

Rank #280 Confirmed

Summary

  • They share 1 benchmark with published results for both. Magistral Medium scores higher in 3 categories and Mistral Small 3.2 in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mistral Small 3.2 leads 18.1 to 8.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 16.2% for Magistral Medium and 40.4% for Mistral Small 3.2.
  • Mistral Small 3.2 is cheaper at $0.0938 / $0.25 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 256K.

Side by side

Magistral Medium and Mistral Small 3.2 specifications
Magistral MediumMistral Small 3.2
ProviderMistral AIMistral AI
Noometry Index35.231.2
Released2025-03-172025-06-20
WeightsOpenOpen
Context window262K256K
Max output16K16K
Input $ / M tokens$2$0.0938
Output $ / M tokens$5$0.25
Results tracked226

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Magistral Medium: 39.1 (#161), Mistral Small 3.2: —

Coding benchmarks
BenchmarkMagistral MediumMistral Small 3.2
SciCode39.2%—
LMArena Coding1319—

Reasoning Mistral Small 3.2 leads

Magistral Medium: 8.6 (#348), Mistral Small 3.2: 18.1 (#287)

Reasoning benchmarks
BenchmarkMagistral MediumMistral Small 3.2
Kagi LLM Benchmark16.2%40.4%
ARC-AGI-20%—
ARC-AGI-16.1%—
CritPt0.3%—
Chess Puzzles—1%
LMArena Hard Prompts1267—
Epoch Capabilities Index—131.74

Math Magistral Medium leads

Magistral Medium: 35.1 (#189), Mistral Small 3.2: 26.3 (#260)

Math benchmarks
BenchmarkMagistral MediumMistral Small 3.2
OTIS Mock AIME 2024-2025—30.3%
LMArena Math1250—

Knowledge Magistral Medium leads

Magistral Medium: 33.5 (#202), Mistral Small 3.2: 26.7 (#256)

Knowledge benchmarks
BenchmarkMagistral MediumMistral Small 3.2
GPQA Diamond—49.1%
LMArena Expert1223—

Multilingual Not comparable

Magistral Medium: 39.6 (#224), Mistral Small 3.2: —

Multilingual benchmarks
BenchmarkMagistral MediumMistral Small 3.2
LMArena Non-English1232—
LMArena Chinese1227—
LMArena French1267—
LMArena German1248—
LMArena Japanese1175—
LMArena Korean1125—
LMArena Russian1224—
LMArena Spanish1271—

Instruction Following Not comparable

Magistral Medium: 66.0 (#211), Mistral Small 3.2: —

Instruction Following benchmarks
BenchmarkMagistral MediumMistral Small 3.2
LMArena Instruction Following1254—

Long Context Not comparable

Magistral Medium: 39.3 (#183), Mistral Small 3.2: —

Long Context benchmarks
BenchmarkMagistral MediumMistral Small 3.2
LMArena Longer Query1295—

Writing & Preference Magistral Medium leads

Magistral Medium: 46.3 (#219), Mistral Small 3.2: 45.0 (#224)

Writing & Preference benchmarks
BenchmarkMagistral MediumMistral Small 3.2
LMArena Text1255—
LMArena Creative Writing1245—
EQ-Bench Creative Writing—1255
LMArena Multi-Turn1275—

Frequently asked questions

Is Magistral Medium better than Mistral Small 3.2?

Magistral Medium is the stronger model overall, scoring 35.2 to 31.2 on the Noometry Index. Mistral Small 3.2 costs 21× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Which is cheaper, Magistral Medium or Mistral Small 3.2?

Mistral Small 3.2 is cheaper. It lists at $0.0938 per million input tokens and $0.25 per million output tokens; Magistral Medium lists at $2 and $5.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 256K.

How many benchmarks do Magistral Medium and Mistral Small 3.2 share?

1 benchmark has published results for both models. Magistral Medium has 22 scored results on Noometry and Mistral Small 3.2 has 6.

Related comparisons

Go deeper