Model comparison

Qwen Max vs QwQ-32B

QwQ-32B is the stronger model overall, scoring 39.8 to 34.7 on the Noometry Index.

Last verified . 21 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Qwen Max scores higher in 1 category and QwQ-32B in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where QwQ-32B leads 38.0 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 16.1% for Qwen Max and 59.2% for QwQ-32B.
  • QwQ-32B has downloadable open weights; the other is API-only.

Side by side

Qwen Max and QwQ-32B specifications
Qwen MaxQwQ-32B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.739.8
Released2024-04-032024-11-28
WeightsProprietaryOpen
Context window33K—
Max output8K—
Input $ / M tokens$1.60—
Output $ / M tokens$6.40—
Results tracked2336

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding QwQ-32B leads

Qwen Max: 30.7 (#292), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkQwen MaxQwQ-32B
Aider Polyglot21.8%20.9%
LMArena Coding12881333
BigCodeBench Instruct—44.6%
LiveBench Coding—72.2%
BigCodeBench Complete—54.4%

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkQwen MaxQwQ-32B
LMArena Hard Prompts12691325
Chess Puzzles—5%
LiveBench Reasoning—83.5%
LiveBench Data Analysis—65%
Epoch Capabilities Index—137.6
ForecastBench—58.3
LiveBench—72%

Math QwQ-32B leads

Qwen Max: 22.3 (#276), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkQwen MaxQwQ-32B
OTIS Mock AIME 2024-202516.1%59.2%
LMArena Math12751359
LiveBench Math—77.8%
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge QwQ-32B leads

Qwen Max: 30.3 (#228), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkQwen MaxQwQ-32B
GPQA Diamond56.1%65.3%
LMArena Expert12481324
Confabulations—15.6%

Multilingual QwQ-32B leads

Qwen Max: 41.8 (#202), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkQwen MaxQwQ-32B
LMArena Non-English12631305
LMArena Chinese12541378
LMArena French13301336
LMArena German12541313
LMArena Japanese12051262
LMArena Korean11421279
LMArena Russian12741297
LMArena Spanish12901354

Instruction Following QwQ-32B leads

Qwen Max: 66.5 (#208), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkQwen MaxQwQ-32B
LMArena Instruction Following12621297
LiveBench Instruction Following—81.8%

Long Context QwQ-32B leads

Qwen Max: 39.4 (#180), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkQwen MaxQwQ-32B
Fiction.LiveBench66.7%83.3%
LMArena Longer Query12881308

Writing & Preference QwQ-32B leads

Qwen Max: 47.8 (#205), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkQwen MaxQwQ-32B
LMArena Text12821329
LMArena Creative Writing12481288
LMArena Multi-Turn12771314
Short-Story Creative Writing—80.2%
EQ-Bench Creative Writing—1257
LiveBench Language—51.4%

Frequently asked questions

Is Qwen Max better than QwQ-32B?

QwQ-32B is the stronger model overall, scoring 39.8 to 34.7 on the Noometry Index.

Is Qwen Max or QwQ-32B better for coding?

QwQ-32B scores higher on coding benchmarks: 35.4 versus 30.7 in the Noometry coding category.

How many benchmarks do Qwen Max and QwQ-32B share?

21 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper