Model comparison

Mistral Medium vs Mistral Small

Mistral Medium is the stronger model overall, scoring 36.3 to 33.4 on the Noometry Index. Mistral Small costs 11× less per token, which makes it the better buy when Mistral Medium's lead doesn't matter for your workload.

Last verified . 29 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Mistral Medium scores higher in 9 categories and Mistral Small in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium leads 28.1 to 16.4.
  • The biggest single-benchmark swing is MATH Level 5: 81.6% for Mistral Medium and 46.8% for Mistral Small.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.

Side by side

Mistral Medium and Mistral Small specifications
Mistral MediumMistral Small
ProviderMistral AIMistral AI
Noometry Index36.333.4
Released2023-12-112024-02-26
WeightsOpenOpen
Context window262K262K
Max output262K256K
Input $ / M tokens$1.50$0.15
Output $ / M tokens$7.50$0.60
Results tracked3639

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Medium: 34.2 (#243), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkMistral MediumMistral Small
SciCode40.2%26.5%
LMArena Coding14341362
ALE-Bench763.98497.62
FrontierCode8%—
WeirdML43.7%—
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%

Agentic & Tool Use Too close to call

Mistral Medium: 28.3 (#90), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkMistral MediumMistral Small
Berkeley Function Calling Leaderboard37.7%37.1%

Reasoning Mistral Medium leads

Mistral Medium: 24.0 (#167), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkMistral MediumMistral Small
Kagi LLM Benchmark50%37.8%
CritPt0%0%
LMArena Hard Prompts14261335
DTBench75.5%70.9%
LMCA26.1%20.6%
LiveBench Reasoning—44.8%
LiveBench Data Analysis—53.7%
Surface Evolver Bench26.9%—
LiveBench—44%

Math Mistral Medium leads

Mistral Medium: 28.1 (#245), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkMistral MediumMistral Small
OTIS Mock AIME 2024-202532.2%5.8%
LMArena Math14081341
MATH Level 581.6%46.8%
ProofBench9%—
LiveBench Math—39.9%
FrontierMath (Feb 2025 set)0.3%—

Knowledge Mistral Small leads

Mistral Medium: 25.0 (#265), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkMistral MediumMistral Small
GPQA Diamond59.5%47.5%
Vectara Hallucination Rate22.7%5.1%
LMArena Expert14081291
Humanity's Last Exam4.5%—
MMLU—68.7%

Multimodal Mistral Medium leads

Mistral Medium: 35.3 (#88), Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkMistral MediumMistral Small
LMArena Vision11721142

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkMistral MediumMistral Small
LMArena Non-English14081315
LMArena Chinese14471340
LMArena French14591337
LMArena German14321340
LMArena Japanese13781275
LMArena Korean13801259
LMArena Russian14111324
LMArena Spanish14331346

Instruction Following Mistral Medium leads

Mistral Medium: 73.7 (#116), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkMistral MediumMistral Small
LMArena Instruction Following13981310
LiveBench Instruction Following—63.7%

Long Context Mistral Medium leads

Mistral Medium: 42.9 (#114), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkMistral MediumMistral Small
LMArena Longer Query14061327

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkMistral MediumMistral Small
LMArena Text14241338
LMArena Creative Writing13911305
LMArena Multi-Turn14181344
Short-Story Creative Writing77.3%—
LiveBench Language—30.5%

Frequently asked questions

Is Mistral Medium better than Mistral Small?

Mistral Medium is the stronger model overall, scoring 36.3 to 33.4 on the Noometry Index. Mistral Small costs 11× less per token, which makes it the better buy when Mistral Medium's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mistral Medium or Mistral Small better for coding?

They score almost the same on coding (34.2 vs 34.0); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Mistral Medium and Mistral Small share?

29 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper