Model comparison

Magistral Medium vs Mistral Small 3.1

Magistral Medium is the stronger model overall, scoring 35.2 to 31.7 on the Noometry Index. Mistral Small 3.1 costs 6.8× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Magistral Medium scores higher in 5 categories and Mistral Small 3.1 in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Medium leads 35.1 to 14.7.
  • Mistral Small 3.1 is cheaper at $0.35 / $0.56 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 128K.

Side by side

Magistral Medium and Mistral Small 3.1 specifications
Magistral MediumMistral Small 3.1
ProviderMistral AIMistral AI
Noometry Index35.231.7
Released2025-03-172025-03-17
WeightsOpenOpen
Context window262K128K
Max output16K102K
Input $ / M tokens$2$0.35
Output $ / M tokens$5$0.56
Results tracked2228

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Magistral Medium: 39.1 (#161), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Coding13191309
SciCode39.2%—

Reasoning Mistral Small 3.1 leads

Magistral Medium: 8.6 (#348), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Hard Prompts12671278
ARC-AGI-20%—
Kagi LLM Benchmark16.2%—
ARC-AGI-16.1%—
CritPt0.3%—
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math Magistral Medium leads

Magistral Medium: 35.1 (#189), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Math12501262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%

Knowledge Magistral Medium leads

Magistral Medium: 33.5 (#202), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Expert12231257
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%

Multimodal Not comparable

Magistral Medium: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Vision—1136

Multilingual Mistral Small 3.1 leads

Magistral Medium: 39.6 (#224), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Non-English12321255
LMArena Chinese12271253
LMArena French12671273
LMArena German12481266
LMArena Japanese11751208
LMArena Korean11251206
LMArena Russian12241263
LMArena Spanish12711283

Instruction Following Magistral Medium leads

Magistral Medium: 66.0 (#211), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Instruction Following12541264
IFEval—75%

Long Context Too close to call

Magistral Medium: 39.3 (#183), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Longer Query12951299

Writing & Preference Magistral Medium leads

Magistral Medium: 46.3 (#219), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkMagistral MediumMistral Small 3.1
LMArena Text12551277
LMArena Creative Writing12451253
LMArena Multi-Turn12751270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is Magistral Medium better than Mistral Small 3.1?

Magistral Medium is the stronger model overall, scoring 35.2 to 31.7 on the Noometry Index. Mistral Small 3.1 costs 6.8× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Which is cheaper, Magistral Medium or Mistral Small 3.1?

Mistral Small 3.1 is cheaper. It lists at $0.35 per million input tokens and $0.56 per million output tokens; Magistral Medium lists at $2 and $5.

Is Magistral Medium or Mistral Small 3.1 better for coding?

They score almost the same on coding (39.1 vs 38.3); test both on your own repository before choosing.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 128K.

How many benchmarks do Magistral Medium and Mistral Small 3.1 share?

17 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper