Model comparison

Qwen3 235B-A22B vs QwQ-32B

Qwen3 235B-A22B is the stronger model overall, scoring 43.5 to 39.8 on the Noometry Index.

Last verified . 27 shared benchmarks.

Qwen3 235B-A22B Alibaba (Qwen)

43.5

Rank #91 Confirmed

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Qwen3 235B-A22B scores higher in 5 categories and QwQ-32B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 235B-A22B leads 49.6 to 37.2.
  • The biggest single-benchmark swing is Aider Polyglot: 59.6% for Qwen3 235B-A22B and 20.9% for QwQ-32B.

Side by side

Qwen3 235B-A22B and QwQ-32B specifications
Qwen3 235B-A22BQwQ-32B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index43.539.8
Released2025-042024-11-28
WeightsOpenOpen
Context window131K—
Max output16K—
Input $ / M tokens$0.70—
Output $ / M tokens$2.80—
Results tracked4936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 235B-A22B leads

Qwen3 235B-A22B: 44.3 (#75), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
Aider Polyglot59.6%20.9%
LMArena Coding14451333
SciCode42.4%—
WeirdML41%—
BigCodeBench Instruct—44.6%
LiveBench Coding—72.2%
BigCodeBench Complete—54.4%

Agentic & Tool Use Not comparable

Qwen3 235B-A22B: 33.9 (#51), QwQ-32B: —

Agentic & Tool Use benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
Berkeley Function Calling Leaderboard52.1%—
Vending-Bench 2-11.34—

Reasoning QwQ-32B leads

Qwen3 235B-A22B: 15.7 (#311), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
Chess Puzzles12%5%
LMArena Hard Prompts14331325
Epoch Capabilities Index143.85137.6
ForecastBench59.758.3
ARC-AGI-21.3%—
SimpleBench31%—
Kagi LLM Benchmark69.4%—
ARC-AGI-111%—
CritPt0%—
LiveBench Reasoning—83.5%
Mystery Game Puzzles9%—
DTBench80.3%—
LiveBench Data Analysis—65%
LMCA29.3%—
LiveBench—72%

Math Qwen3 235B-A22B leads

Qwen3 235B-A22B: 50.4 (#57), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
OTIS Mock AIME 2024-202586.7%59.2%
LMArena Math14321359
Omni-MATH71.8%—
LiveBench Math—77.8%
MATH Level 568.9%—
FrontierMath (Feb 2025 set)8.5%—
FrontierMath Tier 4 (v1)0%—

Knowledge Qwen3 235B-A22B leads

Qwen3 235B-A22B: 49.6 (#73), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
GPQA Diamond80.1%65.3%
Confabulations15.6%15.6%
LMArena Expert14631324
SimpleQA Verified40.4%—
MMLU-Pro84.4%—
Vectara Hallucination Rate9.3%—
GPQA (HELM)72.7%—

Multilingual Qwen3 235B-A22B leads

Qwen3 235B-A22B: 52.3 (#89), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
LMArena Non-English14091305
LMArena Chinese14811378
LMArena French14451336
LMArena German14331313
LMArena Japanese13991262
LMArena Korean13911279
LMArena Russian14111297
LMArena Spanish14301354

Instruction Following Too close to call

Qwen3 235B-A22B: 72.6 (#136), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
LMArena Instruction Following14081297
LiveBench Instruction Following—81.8%
IFEval83.5%—

Long Context QwQ-32B leads

Qwen3 235B-A22B: 46.1 (#26), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
Fiction.LiveBench75%83.3%
LMArena Longer Query14261308

Writing & Preference Qwen3 235B-A22B leads

Qwen3 235B-A22B: 59.6 (#108), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkQwen3 235B-A22BQwQ-32B
LMArena Text14191329
LMArena Creative Writing13841288
Short-Story Creative Writing83%80.2%
EQ-Bench Creative Writing13661257
LMArena Multi-Turn14321314
WildBench86.6%—
LiveBench Language—51.4%

Frequently asked questions

Is Qwen3 235B-A22B better than QwQ-32B?

Qwen3 235B-A22B is the stronger model overall, scoring 43.5 to 39.8 on the Noometry Index.

Is Qwen3 235B-A22B or QwQ-32B better for coding?

Qwen3 235B-A22B scores higher on coding benchmarks: 44.3 versus 35.4 in the Noometry coding category.

How many benchmarks do Qwen3 235B-A22B and QwQ-32B share?

27 benchmarks have published results for both models. Qwen3 235B-A22B has 49 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper