Model comparison

Mistral Medium 3.5 vs Qwen2.5-Coder-32B

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 33.4 on the Noometry Index. Qwen2.5-Coder-32B costs 4.0× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Last verified . 13 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 7 categories and Qwen2.5-Coder-32B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium 3.5 leads 58.5 to 41.6.
  • Qwen2.5-Coder-32B is cheaper at $0.66 / $1 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 33K.

Side by side

Mistral Medium 3.5 and Qwen2.5-Coder-32B specifications
Mistral Medium 3.5Qwen2.5-Coder-32B
ProviderMistral AIAlibaba (Qwen)
Noometry Index40.233.4
Released—2024-09-18
WeightsOpenOpen
Context window262K33K
Max output210K29K
Input $ / M tokens$1.50$0.66
Output $ / M tokens$7.50$1
Results tracked2231

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Mistral Medium 3.5: 36.0 (#213), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Coding14611276
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
LMArena WebDev1264—
BigCodeBench Instruct—49%
LiveBench Coding—56.9%
BigCodeBench Complete—58%
HumanEval+—87.2%
MBPP+—77%

Reasoning Qwen2.5-Coder-32B leads

Mistral Medium 3.5: 17.3 (#295), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Hard Prompts14361251
Epoch Capabilities Index141.35119.49
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
LiveBench Reasoning—42.1%
LiveBench Data Analysis—49.9%
HellaSwag—83%
LiveBench—46.2%
WinoGrande—80.8%

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Math14311251
LiveBench Math—46.6%
GSM8K—93%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Expert14321221
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Qwen2.5-Coder-32B: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Non-English14041205
LMArena Chinese14421222
LMArena Russian13951228
LMArena French1448—
LMArena German1451—
LMArena Korean1385—
LMArena Spanish1409—

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Instruction Following14151223
LiveBench Instruction Following—58.7%

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Longer Query14151251

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Qwen2.5-Coder-32B
LMArena Text14211230
LMArena Creative Writing13741174
LMArena Multi-Turn14231222
EQ-Bench 4993—
LiveBench Language—23.3%

Frequently asked questions

Is Mistral Medium 3.5 better than Qwen2.5-Coder-32B?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 33.4 on the Noometry Index. Qwen2.5-Coder-32B costs 4.0× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.5 or Qwen2.5-Coder-32B?

Qwen2.5-Coder-32B is cheaper. It lists at $0.66 per million input tokens and $1 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Qwen2.5-Coder-32B better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 33K.

How many benchmarks do Mistral Medium 3.5 and Qwen2.5-Coder-32B share?

13 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper