Model comparison

Claude 3.5 Sonnet vs QwQ-32B

QwQ-32B is the stronger model overall, scoring 39.8 to 34.6 on the Noometry Index.

Last verified . 34 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 2 categories and QwQ-32B in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where QwQ-32B leads 38.0 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 59.2% for QwQ-32B.
  • QwQ-32B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and QwQ-32B specifications
Claude 3.5 SonnetQwQ-32B
ProviderAnthropicAlibaba (Qwen)
Noometry Index34.639.8
Released2024-06-202024-11-28
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6036

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
Aider Polyglot51.6%20.9%
BigCodeBench Instruct46.8%44.6%
LiveBench Coding67.1%72.2%
LMArena Coding13421333
BigCodeBench Complete58.6%54.4%
GSO4.6%—
WeirdML40%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), QwQ-32B: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Too close to call

Claude 3.5 Sonnet: 23.1 (#183), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
LiveBench Reasoning56.7%83.5%
LMArena Hard Prompts13051325
LiveBench Data Analysis55%65%
Epoch Capabilities Index133.55137.6
ForecastBench60.758.3
LiveBench59%72%
SimpleBench41.4%—
Chess Puzzles—5%
EnigmaEval0.9%—
DTBench67.8%—

Math QwQ-32B leads

Claude 3.5 Sonnet: 19.2 (#288), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
OTIS Mock AIME 2024-20258.5%59.2%
LiveBench Math52.3%77.8%
LMArena Math13071359
Omni-MATH27.6%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge QwQ-32B leads

Claude 3.5 Sonnet: 28.6 (#245), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
GPQA Diamond55.3%65.3%
Confabulations19.9%15.6%
LMArena Expert12651324
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), QwQ-32B: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual QwQ-32B leads

Claude 3.5 Sonnet: 43.2 (#185), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
LMArena Non-English12831305
LMArena Chinese12721378
LMArena French13051336
LMArena German12971313
LMArena Japanese12341262
LMArena Korean12001279
LMArena Russian13061297
LMArena Spanish12901354

Instruction Following QwQ-32B leads

Claude 3.5 Sonnet: 68.8 (#182), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
LiveBench Instruction Following69.3%81.8%
LMArena Instruction Following12971297
IFEval85.5%—

Long Context QwQ-32B leads

Claude 3.5 Sonnet: 39.9 (#167), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
LMArena Longer Query13111308
Fiction.LiveBench—83.3%

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetQwQ-32B
LMArena Text12981329
LMArena Creative Writing12921288
Short-Story Creative Writing80.3%80.2%
EQ-Bench Creative Writing14511257
LMArena Multi-Turn13261314
LiveBench Language53.8%51.4%
WildBench79.2%—

Frequently asked questions

Is Claude 3.5 Sonnet better than QwQ-32B?

QwQ-32B is the stronger model overall, scoring 39.8 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or QwQ-32B better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 35.4 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and QwQ-32B share?

34 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper