Model comparison

Magistral Medium vs Qwen2.5-VL 72B Instruct

Magistral Medium is the stronger model overall, scoring 35.2 to 29.9 on the Noometry Index.

Last verified . 1 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 1 benchmark with published results for both. Magistral Medium scores higher in 0 categories and Qwen2.5-VL 72B Instruct in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5-VL 72B Instruct leads 20.7 to 8.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 16.2% for Magistral Medium and 36% for Qwen2.5-VL 72B Instruct.
  • Magistral Medium is cheaper at $2 / $5 per million input/output tokens, against $2.80 / $8.40 for Qwen2.5-VL 72B Instruct.
  • Magistral Medium accepts more context: 262K tokens versus 131K.

Side by side

Magistral Medium and Qwen2.5-VL 72B Instruct specifications
Magistral MediumQwen2.5-VL 72B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index35.229.9
Released2025-03-172024-09
WeightsOpenOpen
Context window262K131K
Max output16K8K
Input $ / M tokens$2$2.80
Output $ / M tokens$5$8.40
Results tracked226

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Magistral Medium: 39.1 (#161), Qwen2.5-VL 72B Instruct: —

Coding benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
SciCode39.2%—
LMArena Coding1319—

Agentic & Tool Use Not comparable

Magistral Medium: —, Qwen2.5-VL 72B Instruct: 18.6 (#144)

Agentic & Tool Use benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
OSWorld—5%

Reasoning Qwen2.5-VL 72B Instruct leads

Magistral Medium: 8.6 (#348), Qwen2.5-VL 72B Instruct: 20.7 (#233)

Reasoning benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
Kagi LLM Benchmark16.2%36%
ARC-AGI-20%—
ARC-AGI-16.1%—
CritPt0.3%—
LMArena Hard Prompts1267—

Math Not comparable

Magistral Medium: 35.1 (#189), Qwen2.5-VL 72B Instruct: —

Math benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Math1250—

Knowledge Not comparable

Magistral Medium: 33.5 (#202), Qwen2.5-VL 72B Instruct: —

Knowledge benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Expert1223—

Multimodal Not comparable

Magistral Medium: —, Qwen2.5-VL 72B Instruct: 33.5 (#97)

Multimodal benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Vision—1107
Video-MME—73.5%
GeoBench—62%
SpatialViz-Bench—33.3%

Multilingual Not comparable

Magistral Medium: 39.6 (#224), Qwen2.5-VL 72B Instruct: —

Multilingual benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Non-English1232—
LMArena Chinese1227—
LMArena French1267—
LMArena German1248—
LMArena Japanese1175—
LMArena Korean1125—
LMArena Russian1224—
LMArena Spanish1271—

Instruction Following Not comparable

Magistral Medium: 66.0 (#211), Qwen2.5-VL 72B Instruct: —

Instruction Following benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Instruction Following1254—

Long Context Not comparable

Magistral Medium: 39.3 (#183), Qwen2.5-VL 72B Instruct: —

Long Context benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Longer Query1295—

Writing & Preference Not comparable

Magistral Medium: 46.3 (#219), Qwen2.5-VL 72B Instruct: —

Writing & Preference benchmarks
BenchmarkMagistral MediumQwen2.5-VL 72B Instruct
LMArena Text1255—
LMArena Creative Writing1245—
LMArena Multi-Turn1275—

Frequently asked questions

Is Magistral Medium better than Qwen2.5-VL 72B Instruct?

Magistral Medium is the stronger model overall, scoring 35.2 to 29.9 on the Noometry Index.

Which is cheaper, Magistral Medium or Qwen2.5-VL 72B Instruct?

Magistral Medium is cheaper. It lists at $2 per million input tokens and $5 per million output tokens; Qwen2.5-VL 72B Instruct lists at $2.80 and $8.40.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 131K.

How many benchmarks do Magistral Medium and Qwen2.5-VL 72B Instruct share?

1 benchmark has published results for both models. Magistral Medium has 22 scored results on Noometry and Qwen2.5-VL 72B Instruct has 6.

Related comparisons

Go deeper