Model comparison

Mistral Medium 3.5 vs Mistral Small

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 33.4 on the Noometry Index. Mistral Small costs 11× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 8 categories and Mistral Small in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium 3.5 leads 39.1 to 16.4.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.

Side by side

Mistral Medium 3.5 and Mistral Small specifications
Mistral Medium 3.5Mistral Small
ProviderMistral AIMistral AI
Noometry Index40.233.4
Released—2024-02-26
WeightsOpenOpen
Context window262K262K
Max output210K256K
Input $ / M tokens$1.50$0.15
Output $ / M tokens$7.50$0.60
Results tracked2239

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Mistral Medium 3.5: 36.0 (#213), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Coding14611362
LMArena WebDev1264—
SciCode—26.5%
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%
ALE-Bench—497.62

Agentic & Tool Use Not comparable

Mistral Medium 3.5: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.5Mistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Mistral Small leads

Mistral Medium 3.5: 17.3 (#295), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Mistral Small
Kagi LLM Benchmark41.4%37.8%
LMArena Hard Prompts14361335
NYT Connections (extended)12.9%—
CritPt—0%
LiveBench Reasoning—44.8%
DTBench—70.9%
LiveBench Data Analysis—53.7%
LMCA—20.6%
Epoch Capabilities Index141.35—
LiveBench—44%

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Math14311341
OTIS Mock AIME 2024-2025—5.8%
LiveBench Math—39.9%
MATH Level 5—46.8%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Expert14321291
GPQA Diamond—47.5%
Vectara Hallucination Rate—5.1%
MMLU—68.7%

Multimodal Mistral Medium 3.5 leads

Mistral Medium 3.5: 38.3 (#65), Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Vision12231142

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Non-English14041315
LMArena Chinese14421340
LMArena French14481337
LMArena German14511340
LMArena Korean13851259
LMArena Russian13951324
LMArena Spanish14091346
LMArena Japanese—1275

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Instruction Following14151310
LiveBench Instruction Following—63.7%

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Longer Query14151327

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Mistral Small
LMArena Text14211338
LMArena Creative Writing13741305
LMArena Multi-Turn14231344
EQ-Bench 4993—
LiveBench Language—30.5%

Frequently asked questions

Is Mistral Medium 3.5 better than Mistral Small?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 33.4 on the Noometry Index. Mistral Small costs 11× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.5 or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Mistral Small better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Mistral Medium 3.5 and Mistral Small share?

18 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper