Model comparison

Magistral Medium vs Mistral Large

Magistral Medium is the stronger model overall, scoring 35.2 to 31.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Magistral Medium scores higher in 5 categories and Mistral Large in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Medium leads 35.1 to 18.2.
  • Magistral Medium is cheaper at $2 / $5 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Magistral Medium accepts more context: 262K tokens versus 131K.

Side by side

Magistral Medium and Mistral Large specifications
Magistral MediumMistral Large
ProviderMistral AIMistral AI
Noometry Index35.231.9
Released2025-03-172024-02-26
WeightsOpenOpen
Context window262K131K
Max output16K16K
Input $ / M tokens$2$2
Output $ / M tokens$5$6
Results tracked2251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Magistral Medium: 39.1 (#161), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkMagistral MediumMistral Large
SciCode39.2%36.2%
LMArena Coding13191277
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Magistral Medium: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkMagistral MediumMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Mistral Large leads

Magistral Medium: 8.6 (#348), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkMagistral MediumMistral Large
CritPt0.3%0%
LMArena Hard Prompts12671257
ARC-AGI-20%—
SimpleBench—22.5%
Kagi LLM Benchmark16.2%—
ARC-AGI-16.1%—
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Magistral Medium leads

Magistral Medium: 35.1 (#189), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkMagistral MediumMistral Large
LMArena Math12501262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Magistral Medium leads

Magistral Medium: 33.5 (#202), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkMagistral MediumMistral Large
LMArena Expert12231232
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Too close to call

Magistral Medium: 39.6 (#224), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkMagistral MediumMistral Large
LMArena Non-English12321237
LMArena Chinese12271240
LMArena French12671325
LMArena German12481254
LMArena Japanese11751188
LMArena Korean11251202
LMArena Russian12241257
LMArena Spanish12711268

Instruction Following Mistral Large leads

Magistral Medium: 66.0 (#211), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkMagistral MediumMistral Large
LMArena Instruction Following12541249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Magistral Medium leads

Magistral Medium: 39.3 (#183), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkMagistral MediumMistral Large
LMArena Longer Query12951261

Writing & Preference Magistral Medium leads

Magistral Medium: 46.3 (#219), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkMagistral MediumMistral Large
LMArena Text12551266
LMArena Creative Writing12451243
LMArena Multi-Turn12751260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Magistral Medium better than Mistral Large?

Magistral Medium is the stronger model overall, scoring 35.2 to 31.9 on the Noometry Index.

Which is cheaper, Magistral Medium or Mistral Large?

Magistral Medium is cheaper. It lists at $2 per million input tokens and $5 per million output tokens; Mistral Large lists at $2 and $6.

Is Magistral Medium or Mistral Large better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 131K.

How many benchmarks do Magistral Medium and Mistral Large share?

19 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper