Model comparison

Mistral vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 29.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral scores higher in 0 categories and Mistral Medium in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 37.0.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mistral and Mistral Medium specifications
MistralMistral Medium
ProviderMistral AIMistral AI
Noometry Index29.936.3
Released—2023-12-11
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked2236

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral: 33.8 (#250), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkMistralMistral Medium
LMArena Coding11621434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Mistral: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkMistralMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Mistral: 22.2 (#200), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkMistralMistral Medium
LMArena Hard Prompts11491426
Kagi LLM Benchmark—50%
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Mistral Medium leads

Mistral: 22.3 (#278), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkMistralMistral Medium
LMArena Math11801408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
Omni-MATH7.2%—
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Mistral Medium leads

Mistral: 16.6 (#288), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkMistralMistral Medium
LMArena Expert11251408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
MMLU-Pro27.7%—
Vectara Hallucination Rate—22.7%
GPQA (HELM)30.3%—

Multimodal Not comparable

Mistral: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkMistralMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Mistral: 32.8 (#254), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkMistralMistral Medium
LMArena Non-English11291408
LMArena Chinese11091447
LMArena French11801459
LMArena German11551432
LMArena Japanese10131378
LMArena Korean10321380
LMArena Russian11681411
LMArena Spanish11431433

Instruction Following Mistral Medium leads

Mistral: 52.6 (#288), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkMistralMistral Medium
LMArena Instruction Following11521398
IFEval56.8%—

Long Context Mistral Medium leads

Mistral: 35.0 (#245), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkMistralMistral Medium
LMArena Longer Query11531406

Writing & Preference Mistral Medium leads

Mistral: 37.0 (#260), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkMistralMistral Medium
LMArena Text11651424
LMArena Creative Writing11581391
LMArena Multi-Turn11471418
Short-Story Creative Writing—77.3%
WildBench66%—

Frequently asked questions

Is Mistral better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 29.9 on the Noometry Index.

Is Mistral or Mistral Medium better for coding?

They score almost the same on coding (33.8 vs 34.2); test both on your own repository before choosing.

How many benchmarks do Mistral and Mistral Medium share?

17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper