Model comparison

Magistral Small vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 30.2 on the Noometry Index. Magistral Small costs 4.0× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 6 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Magistral Small scores higher in 3 categories and Mistral Large in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mistral Large leads 15.8 to 6.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 30% for Magistral Small and 8.5% for Mistral Large.
  • Magistral Small is cheaper at $0.50 / $1.50 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 128K.

Side by side

Magistral Small and Mistral Large specifications
Magistral SmallMistral Large
ProviderMistral AIMistral AI
Noometry Index30.231.9
Released2025-06-102024-02-26
WeightsOpenOpen
Context window128K131K
Max output40K16K
Input $ / M tokens$0.50$2
Output $ / M tokens$1.50$6
Results tracked1051

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Magistral Small: 38.4 (#176), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkMagistral SmallMistral Large
SciCode35.2%36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Magistral Small: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkMagistral SmallMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Mistral Large leads

Magistral Small: 6.8 (#350), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkMagistral SmallMistral Large
CritPt0.3%0%
DTBench61.3%65.1%
Epoch Capabilities Index133.19128.52
ARC-AGI-20%—
SimpleBench—22.5%
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
Chess Puzzles3%—
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
LiveBench Data Analysis—50.1%
LMCA—16.7%
ForecastBench—57.1
LiveBench—48.4%

Math Magistral Small leads

Magistral Small: 26.2 (#261), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkMagistral SmallMistral Large
OTIS Mock AIME 2024-202530%8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Too close to call

Magistral Small: 30.9 (#223), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkMagistral SmallMistral Large
GPQA Diamond56.1%51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Magistral Small: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkMagistral SmallMistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Not comparable

Magistral Small: —, Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkMagistral SmallMistral Large
LiveBench Instruction Following—67.9%
IFEval—87.7%
LMArena Instruction Following—1249

Long Context Not comparable

Magistral Small: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkMagistral SmallMistral Large
LMArena Longer Query—1261

Writing & Preference Not comparable

Magistral Small: —, Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkMagistral SmallMistral Large
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Magistral Small better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 30.2 on the Noometry Index. Magistral Small costs 4.0× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Magistral Small or Mistral Large?

Magistral Small is cheaper. It lists at $0.50 per million input tokens and $1.50 per million output tokens; Mistral Large lists at $2 and $6.

Is Magistral Small or Mistral Large better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 128K.

How many benchmarks do Magistral Small and Mistral Large share?

6 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper