Model comparison

Mistral Medium 3.5 vs Qwen Max

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 34.7 on the Noometry Index.

Last verified . 16 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 7 categories and Qwen Max in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium 3.5 leads 39.1 to 22.3.
  • Qwen Max is cheaper at $1.60 / $6.40 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 33K.
  • Mistral Medium 3.5 has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.5 and Qwen Max specifications
Mistral Medium 3.5Qwen Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index40.234.7
Released—2024-04-03
WeightsOpenProprietary
Context window262K33K
Max output210K8K
Input $ / M tokens$1.50$1.60
Output $ / M tokens$7.50$6.40
Results tracked2223

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Mistral Medium 3.5: 36.0 (#213), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Coding14611288
Aider Polyglot—21.8%
LMArena WebDev1264—

Reasoning Qwen Max leads

Mistral Medium 3.5: 17.3 (#295), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Hard Prompts14361269
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
Epoch Capabilities Index141.35—

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Math14311275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Expert14321248
GPQA Diamond—56.1%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Qwen Max: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Non-English14041263
LMArena Chinese14421254
LMArena French14481330
LMArena German14511254
LMArena Korean13851142
LMArena Russian13951274
LMArena Spanish14091290
LMArena Japanese—1205

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Instruction Following14151262

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Longer Query14151288
Fiction.LiveBench—66.7%

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Qwen Max
LMArena Text14211282
LMArena Creative Writing13741248
LMArena Multi-Turn14231277
EQ-Bench 4993—

Frequently asked questions

Is Mistral Medium 3.5 better than Qwen Max?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 34.7 on the Noometry Index.

Which is cheaper, Mistral Medium 3.5 or Qwen Max?

Qwen Max is cheaper. It lists at $1.60 per million input tokens and $6.40 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Qwen Max better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 33K.

How many benchmarks do Mistral Medium 3.5 and Qwen Max share?

16 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper