Model comparison

Mistral Medium 3.1 vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 31.9 on the Noometry Index. Mistral Medium 3.1 costs 1.5× less per token, which makes it the better buy when Qwen3 32B's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3 32B leads 20.2 to 10.6.
  • Mistral Medium 3.1 is cheaper at $0.40 / $2 per million input/output tokens, against $0.70 / $2.80 for Qwen3 32B.
  • Qwen3 32B has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.1 and Qwen3 32B specifications
Mistral Medium 3.1Qwen3 32B
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.939.2
Released—2025-04
WeightsProprietaryOpen
Context window131K131K
Max output105K16K
Input $ / M tokens$0.40$0.70
Output $ / M tokens$2$2.80
Results tracked326

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
Aider Polyglot—40%
SciCode—35.4%
LMArena Coding—1358

Agentic & Tool Use Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
Berkeley Function Calling Leaderboard—48.7%

Reasoning Qwen3 32B leads

Mistral Medium 3.1: 10.6 (#341), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
Kagi LLM Benchmark—54.9%
NYT Connections (extended)6.5%—
CritPt—0.3%
Chess Puzzles—5%
Thematic Generalization20.3%—
LMArena Hard Prompts—1334
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
OTIS Mock AIME 2024-2025—66.9%
LMArena Math—1399

Knowledge Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%
LMArena Expert—1362

Multilingual Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
LMArena Non-English—1317
LMArena Chinese—1357
LMArena German—1341
LMArena Russian—1311

Instruction Following Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
LMArena Instruction Following—1305

Long Context Not comparable

Mistral Medium 3.1: —, Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
Fiction.LiveBench—74.2%
LMArena Longer Query—1327

Writing & Preference Mistral Medium 3.1 leads

Mistral Medium 3.1: 55.5 (#145), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1Qwen3 32B
LMArena Text—1340
LMArena Creative Writing—1297
EQ-Bench Creative Writing1476—
LMArena Multi-Turn—1331

Frequently asked questions

Is Mistral Medium 3.1 better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 31.9 on the Noometry Index. Mistral Medium 3.1 costs 1.5× less per token, which makes it the better buy when Qwen3 32B's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.1 or Qwen3 32B?

Mistral Medium 3.1 is cheaper. It lists at $0.40 per million input tokens and $2 per million output tokens; Qwen3 32B lists at $0.70 and $2.80.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Mistral Medium 3.1 and Qwen3 32B share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper