Model comparison

Mistral Medium 3.5 vs Mistral Small 3

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 31.2 on the Noometry Index. Mistral Small 3 costs 52× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Last verified . 16 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Mistral Small 3 Mistral AI

31.2

Rank #278 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 6 categories and Mistral Small 3 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium 3.5 leads 58.5 to 32.2.
  • Mistral Small 3 is cheaper at $0.05 / $0.08 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 33K.

Side by side

Mistral Medium 3.5 and Mistral Small 3 specifications
Mistral Medium 3.5Mistral Small 3
ProviderMistral AIMistral AI
Noometry Index40.231.2
Released—2025-01-30
WeightsOpenOpen
Context window262K33K
Max output210K16K
Input $ / M tokens$1.50$0.05
Output $ / M tokens$7.50$0.08
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Medium 3.5: 36.0 (#213), Mistral Small 3: 36.5 (#207)

Coding benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Coding14611246
LMArena WebDev1264—
BigCodeBench Instruct—45.3%
BigCodeBench Complete—50.4%

Reasoning Mistral Small 3 leads

Mistral Medium 3.5: 17.3 (#295), Mistral Small 3: 18.9 (#273)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Hard Prompts14361233
Epoch Capabilities Index141.35127.07
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
Chess Puzzles—0%

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Mistral Small 3: 16.3 (#295)

Math benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Math14311240
OTIS Mock AIME 2024-2025—6.7%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Mistral Small 3: 25.1 (#263)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Expert14321202
GPQA Diamond—47.3%
Confabulations—25.2%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Mistral Small 3: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Mistral Small 3: 37.3 (#236)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Non-English14041198
LMArena Chinese14421204
LMArena French14481203
LMArena German14511211
LMArena Korean13851188
LMArena Russian13951216
LMArena Japanese—1111
LMArena Spanish1409—

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Mistral Small 3: 63.7 (#229)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Instruction Following14151214

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Mistral Small 3: 37.8 (#211)

Long Context benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Longer Query14151246

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Mistral Small 3: 32.2 (#280)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Mistral Small 3
LMArena Text14211234
LMArena Creative Writing13741195
LMArena Multi-Turn14231217
EQ-Bench Creative Writing—707
EQ-Bench 4993—

Frequently asked questions

Is Mistral Medium 3.5 better than Mistral Small 3?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 31.2 on the Noometry Index. Mistral Small 3 costs 52× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.5 or Mistral Small 3?

Mistral Small 3 is cheaper. It lists at $0.05 per million input tokens and $0.08 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Mistral Small 3 better for coding?

They score almost the same on coding (36.0 vs 36.5); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 33K.

How many benchmarks do Mistral Medium 3.5 and Mistral Small 3 share?

16 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Mistral Small 3 has 24.

Related comparisons

Go deeper