Model comparison

Qwen2.5-Coder-32B vs QwQ-32B

QwQ-32B is the stronger model overall, scoring 39.8 to 33.4 on the Noometry Index.

Last verified . 23 shared benchmarks.

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Qwen2.5-Coder-32B scores higher in 0 categories and QwQ-32B in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where QwQ-32B leads 35.4 to 22.6.
  • The biggest single-benchmark swing is LiveBench Reasoning: 42.1% for Qwen2.5-Coder-32B and 83.5% for QwQ-32B.

Side by side

Qwen2.5-Coder-32B and QwQ-32B specifications
Qwen2.5-Coder-32BQwQ-32B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index33.439.8
Released2024-09-182024-11-28
WeightsOpenOpen
Context window33K—
Max output29K—
Input $ / M tokens$0.66—
Output $ / M tokens$1—
Results tracked3136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding QwQ-32B leads

Qwen2.5-Coder-32B: 22.6 (#333), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
Aider Polyglot16.4%20.9%
BigCodeBench Instruct49%44.6%
LiveBench Coding56.9%72.2%
LMArena Coding12761333
BigCodeBench Complete58%54.4%
SWE-bench Verified (bash only)9%—
HumanEval+87.2%—
MBPP+77%—

Reasoning QwQ-32B leads

Qwen2.5-Coder-32B: 21.2 (#225), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LiveBench Reasoning42.1%83.5%
LMArena Hard Prompts12511325
LiveBench Data Analysis49.9%65%
Epoch Capabilities Index119.49137.6
LiveBench46.2%72%
Chess Puzzles—5%
ForecastBench—58.3
HellaSwag83%—
WinoGrande80.8%—

Math QwQ-32B leads

Qwen2.5-Coder-32B: 33.3 (#204), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LiveBench Math46.6%77.8%
LMArena Math12511359
OTIS Mock AIME 2024-2025—59.2%
GSM8K93%—

Knowledge QwQ-32B leads

Qwen2.5-Coder-32B: 33.4 (#203), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LMArena Expert12211324
GPQA Diamond—65.3%
Confabulations—15.6%
ARC (AI2) Challenge70.5%—
MMLU79.1%—

Multilingual QwQ-32B leads

Qwen2.5-Coder-32B: 37.8 (#235), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LMArena Non-English12051305
LMArena Chinese12221378
LMArena Russian12281297
LMArena French—1336
LMArena German—1313
LMArena Japanese—1262
LMArena Korean—1279
LMArena Spanish—1354

Instruction Following QwQ-32B leads

Qwen2.5-Coder-32B: 61.4 (#245), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LiveBench Instruction Following58.7%81.8%
LMArena Instruction Following12231297

Long Context QwQ-32B leads

Qwen2.5-Coder-32B: 38.0 (#208), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LMArena Longer Query12511308
Fiction.LiveBench—83.3%

Writing & Preference QwQ-32B leads

Qwen2.5-Coder-32B: 41.6 (#240), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkQwen2.5-Coder-32BQwQ-32B
LMArena Text12301329
LMArena Creative Writing11741288
LMArena Multi-Turn12221314
LiveBench Language23.3%51.4%
Short-Story Creative Writing—80.2%
EQ-Bench Creative Writing—1257

Frequently asked questions

Is Qwen2.5-Coder-32B better than QwQ-32B?

QwQ-32B is the stronger model overall, scoring 39.8 to 33.4 on the Noometry Index.

Is Qwen2.5-Coder-32B or QwQ-32B better for coding?

QwQ-32B scores higher on coding benchmarks: 35.4 versus 22.6 in the Noometry coding category.

How many benchmarks do Qwen2.5-Coder-32B and QwQ-32B share?

23 benchmarks have published results for both models. Qwen2.5-Coder-32B has 31 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper