Model comparison

Mistral vs Qwen2.5 72B Instruct

Qwen2.5 72B Instruct is the stronger model overall, scoring 31.9 to 29.9 on the Noometry Index.

Last verified . 22 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Mistral scores higher in 2 categories and Qwen2.5 72B Instruct in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Qwen2.5 72B Instruct leads 65.5 to 52.6.
  • The biggest single-benchmark swing is MMLU-Pro: 27.7% for Mistral and 63.1% for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct has downloadable open weights; the other is API-only.

Side by side

Mistral and Qwen2.5 72B Instruct specifications
MistralQwen2.5 72B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.931.9
Released—2024-09
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$1.40
Output $ / M tokens—$5.60
Results tracked2243

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral: 33.8 (#250), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkMistralQwen2.5 72B Instruct
LMArena Coding11621292
WeirdML—16%
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%

Agentic & Tool Use Not comparable

Mistral: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkMistralQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Too close to call

Mistral: 22.2 (#200), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkMistralQwen2.5 72B Instruct
LMArena Hard Prompts11491271
DTBench—62.9%
LMCA—13.4%
BIG-Bench Hard—79.8%
Epoch Capabilities Index—129
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Mistral leads

Mistral: 22.3 (#278), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkMistralQwen2.5 72B Instruct
Omni-MATH7.2%33%
LMArena Math11801283
OTIS Mock AIME 2024-2025—8.1%
MATH Level 5—63.2%

Knowledge Qwen2.5 72B Instruct leads

Mistral: 16.6 (#288), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkMistralQwen2.5 72B Instruct
MMLU-Pro27.7%63.1%
GPQA (HELM)30.3%42.6%
LMArena Expert11251245
GPQA Diamond—49.1%
Confabulations—19.1%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multilingual Qwen2.5 72B Instruct leads

Mistral: 32.8 (#254), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkMistralQwen2.5 72B Instruct
LMArena Non-English11291252
LMArena Chinese11091272
LMArena French11801280
LMArena German11551234
LMArena Japanese10131180
LMArena Korean10321188
LMArena Russian11681264
LMArena Spanish11431256

Instruction Following Qwen2.5 72B Instruct leads

Mistral: 52.6 (#288), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkMistralQwen2.5 72B Instruct
IFEval56.8%80.6%
LMArena Instruction Following11521254

Long Context Qwen2.5 72B Instruct leads

Mistral: 35.0 (#245), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkMistralQwen2.5 72B Instruct
LMArena Longer Query11531282

Writing & Preference Qwen2.5 72B Instruct leads

Mistral: 37.0 (#260), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkMistralQwen2.5 72B Instruct
LMArena Text11651269
LMArena Creative Writing11581221
WildBench66%80.2%
LMArena Multi-Turn11471272

Frequently asked questions

Is Mistral better than Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is the stronger model overall, scoring 31.9 to 29.9 on the Noometry Index.

Is Mistral or Qwen2.5 72B Instruct better for coding?

They score almost the same on coding (33.8 vs 33.2); test both on your own repository before choosing.

How many benchmarks do Mistral and Qwen2.5 72B Instruct share?

22 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper