Model comparison

Mistral Medium vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 36.3 on the Noometry Index.

Last verified . 22 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Mistral Medium scores higher in 4 categories and Qwen3 32B in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 32B leads 40.0 to 25.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 32.2% for Mistral Medium and 66.9% for Qwen3 32B.
  • Qwen3 32B is cheaper at $0.70 / $2.80 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 131K.

Side by side

Mistral Medium and Qwen3 32B specifications
Mistral MediumQwen3 32B
ProviderMistral AIAlibaba (Qwen)
Noometry Index36.339.2
Released2023-12-112025-04
WeightsOpenOpen
Context window262K131K
Max output262K16K
Input $ / M tokens$1.50$0.70
Output $ / M tokens$7.50$2.80
Results tracked3626

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 32B leads

Mistral Medium: 34.2 (#243), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkMistral MediumQwen3 32B
SciCode40.2%35.4%
LMArena Coding14341358
FrontierCode8%—
Aider Polyglot—40%
WeirdML43.7%—
ALE-Bench763.98—

Agentic & Tool Use Qwen3 32B leads

Mistral Medium: 28.3 (#90), Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkMistral MediumQwen3 32B
Berkeley Function Calling Leaderboard37.7%48.7%

Reasoning Mistral Medium leads

Mistral Medium: 24.0 (#167), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkMistral MediumQwen3 32B
Kagi LLM Benchmark50%54.9%
CritPt0%0.3%
LMArena Hard Prompts14261334
DTBench75.5%67.5%
LMCA26.1%17.3%
Chess Puzzles—5%
Surface Evolver Bench26.9%—
Epoch Capabilities Index—138.51

Math Qwen3 32B leads

Mistral Medium: 28.1 (#245), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkMistral MediumQwen3 32B
OTIS Mock AIME 2024-202532.2%66.9%
LMArena Math14081399
ProofBench9%—
MATH Level 581.6%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen3 32B leads

Mistral Medium: 25.0 (#265), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkMistral MediumQwen3 32B
GPQA Diamond59.5%65.7%
Vectara Hallucination Rate22.7%5.9%
LMArena Expert14081362
Humanity's Last Exam4.5%—

Multimodal Not comparable

Mistral Medium: 35.3 (#88), Qwen3 32B: —

Multimodal benchmarks
BenchmarkMistral MediumQwen3 32B
LMArena Vision1172—

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkMistral MediumQwen3 32B
LMArena Non-English14081317
LMArena Chinese14471357
LMArena German14321341
LMArena Russian14111311
LMArena French1459—
LMArena Japanese1378—
LMArena Korean1380—
LMArena Spanish1433—

Instruction Following Mistral Medium leads

Mistral Medium: 73.7 (#116), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkMistral MediumQwen3 32B
LMArena Instruction Following13981305

Long Context Too close to call

Mistral Medium: 42.9 (#114), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkMistral MediumQwen3 32B
LMArena Longer Query14061327
Fiction.LiveBench—74.2%

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkMistral MediumQwen3 32B
LMArena Text14241340
LMArena Creative Writing13911297
LMArena Multi-Turn14181331
Short-Story Creative Writing77.3%—

Frequently asked questions

Is Mistral Medium better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 36.3 on the Noometry Index.

Which is cheaper, Mistral Medium or Qwen3 32B?

Qwen3 32B is cheaper. It lists at $0.70 per million input tokens and $2.80 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mistral Medium or Qwen3 32B better for coding?

Qwen3 32B scores higher on coding benchmarks: 37.7 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 131K.

How many benchmarks do Mistral Medium and Qwen3 32B share?

22 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper