Model comparison

Mistral Medium vs Qwen2.5 72B Instruct

Mistral Medium is the stronger model overall, scoring 36.3 to 31.9 on the Noometry Index.

Last verified . 23 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Mistral Medium scores higher in 8 categories and Qwen2.5 72B Instruct in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 46.7.
  • The biggest single-benchmark swing is WeirdML: 43.7% for Mistral Medium and 16% for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct is cheaper at $1.40 / $5.60 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 131K.

Side by side

Mistral Medium and Qwen2.5 72B Instruct specifications
Mistral MediumQwen2.5 72B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index36.331.9
Released2023-12-112024-09
WeightsOpenOpen
Context window262K131K
Max output262K8K
Input $ / M tokens$1.50$1.40
Output $ / M tokens$7.50$5.60
Results tracked3643

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Medium: 34.2 (#243), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
WeirdML43.7%16%
LMArena Coding14341292
FrontierCode8%—
SciCode40.2%—
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%
ALE-Bench763.98—

Agentic & Tool Use Mistral Medium leads

Mistral Medium: 28.3 (#90), Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
Berkeley Function Calling Leaderboard37.7%—
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Mistral Medium leads

Mistral Medium: 24.0 (#167), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
LMArena Hard Prompts14261271
DTBench75.5%62.9%
LMCA26.1%13.4%
Kagi LLM Benchmark50%—
CritPt0%—
Surface Evolver Bench26.9%—
BIG-Bench Hard—79.8%
Epoch Capabilities Index—129
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Mistral Medium leads

Mistral Medium: 28.1 (#245), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
OTIS Mock AIME 2024-202532.2%8.1%
LMArena Math14081283
MATH Level 581.6%63.2%
ProofBench9%—
Omni-MATH—33%
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen2.5 72B Instruct leads

Mistral Medium: 25.0 (#265), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
GPQA Diamond59.5%49.1%
LMArena Expert14081245
Humanity's Last Exam4.5%—
MMLU-Pro—63.1%
Confabulations—19.1%
Vectara Hallucination Rate22.7%—
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multimodal Not comparable

Mistral Medium: 35.3 (#88), Qwen2.5 72B Instruct: —

Multimodal benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
LMArena Vision1172—

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
LMArena Non-English14081252
LMArena Chinese14471272
LMArena French14591280
LMArena German14321234
LMArena Japanese13781180
LMArena Korean13801188
LMArena Russian14111264
LMArena Spanish14331256

Instruction Following Mistral Medium leads

Mistral Medium: 73.7 (#116), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
LMArena Instruction Following13981254
IFEval—80.6%

Long Context Mistral Medium leads

Mistral Medium: 42.9 (#114), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
LMArena Longer Query14061282

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkMistral MediumQwen2.5 72B Instruct
LMArena Text14241269
LMArena Creative Writing13911221
LMArena Multi-Turn14181272
Short-Story Creative Writing77.3%—
WildBench—80.2%

Frequently asked questions

Is Mistral Medium better than Qwen2.5 72B Instruct?

Mistral Medium is the stronger model overall, scoring 36.3 to 31.9 on the Noometry Index.

Which is cheaper, Mistral Medium or Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is cheaper. It lists at $1.40 per million input tokens and $5.60 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mistral Medium or Qwen2.5 72B Instruct better for coding?

They score almost the same on coding (34.2 vs 33.2); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 131K.

How many benchmarks do Mistral Medium and Qwen2.5 72B Instruct share?

23 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper