Model comparison

Mistral Medium 3.5 vs Qwen2.5-Max

Mistral Medium 3.5 and Qwen2.5-Max score almost the same on the Noometry Index (40.2 vs 40.7), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Qwen2.5-Max Alibaba (Qwen)

40.7

Rank #146 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 6 categories and Qwen2.5-Max in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5-Max leads 25.6 to 17.3.
  • Mistral Medium 3.5 has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.5 and Qwen2.5-Max specifications
Mistral Medium 3.5Qwen2.5-Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index40.240.7
Released—2025-01-25
WeightsOpenProprietary
Context window262K—
Max output210K—
Input $ / M tokens$1.50—
Output $ / M tokens$7.50—
Results tracked2227

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5-Max leads

Mistral Medium 3.5: 36.0 (#213), Qwen2.5-Max: 41.8 (#117)

Coding benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Coding14611359
LMArena WebDev1264—
LiveBench Coding—64.4%

Reasoning Qwen2.5-Max leads

Mistral Medium 3.5: 17.3 (#295), Qwen2.5-Max: 25.6 (#147)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Hard Prompts14361360
Epoch Capabilities Index141.35132.53
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
LiveBench Reasoning—51.4%
LiveBench Data Analysis—67.9%
LiveBench—62.3%

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Qwen2.5-Max: 36.9 (#162)

Math benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Math14311369
LiveBench Math—58.4%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Qwen2.5-Max: 35.3 (#186)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Expert14321337
Confabulations—21.8%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Qwen2.5-Max: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Qwen2.5-Max: 48.1 (#146)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Non-English14041352
LMArena Chinese14421382
LMArena French14481396
LMArena German14511350
LMArena Korean13851304
LMArena Russian13951353
LMArena Spanish14091377
LMArena Japanese—1300

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Qwen2.5-Max: 71.3 (#152)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Instruction Following14151335
LiveBench Instruction Following—75.3%

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Qwen2.5-Max: 41.4 (#142)

Long Context benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Longer Query14151358

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Qwen2.5-Max: 55.4 (#146)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Max
LMArena Text14211367
LMArena Creative Writing13741339
LMArena Multi-Turn14231364
Short-Story Creative Writing—72.9%
EQ-Bench 4993—
LiveBench Language—56.3%

Frequently asked questions

Is Mistral Medium 3.5 better than Qwen2.5-Max?

Mistral Medium 3.5 and Qwen2.5-Max score almost the same on the Noometry Index (40.2 vs 40.7), so choose on price, context window or the category you care about most.

Is Mistral Medium 3.5 or Qwen2.5-Max better for coding?

Qwen2.5-Max scores higher on coding benchmarks: 41.8 versus 36.0 in the Noometry coding category.

How many benchmarks do Mistral Medium 3.5 and Qwen2.5-Max share?

17 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Qwen2.5-Max has 27.

Related comparisons

Go deeper