Model comparison

Claude 3.5 Sonnet vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 34.6 on the Noometry Index.

Last verified . 18 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 2 categories and Qwen3 32B in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3 32B leads 39.7 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 66.9% for Qwen3 32B.
  • Qwen3 32B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Qwen3 32B specifications
Claude 3.5 SonnetQwen3 32B
ProviderAnthropicAlibaba (Qwen)
Noometry Index34.639.2
Released2024-06-202025-04
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked6026

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
Aider Polyglot51.6%40%
LMArena Coding13421358
SciCode—35.4%
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Too close to call

Claude 3.5 Sonnet: 32.3 (#67), Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
Berkeley Function Calling Leaderboard—48.7%
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
LMArena Hard Prompts13051334
DTBench67.8%67.5%
Epoch Capabilities Index133.55138.51
SimpleBench41.4%—
Kagi LLM Benchmark—54.9%
CritPt—0.3%
Chess Puzzles—5%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LiveBench Data Analysis55%—
LMCA—17.3%
ForecastBench60.7—
LiveBench59%—

Math Qwen3 32B leads

Claude 3.5 Sonnet: 19.2 (#288), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
OTIS Mock AIME 2024-20258.5%66.9%
LMArena Math13071399
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Qwen3 32B leads

Claude 3.5 Sonnet: 28.6 (#245), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
GPQA Diamond55.3%65.7%
LMArena Expert12651362
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
Vectara Hallucination Rate—5.9%
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Qwen3 32B: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Qwen3 32B leads

Claude 3.5 Sonnet: 43.2 (#185), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
LMArena Non-English12831317
LMArena Chinese12721357
LMArena German12971341
LMArena Russian13061311
LMArena French1305—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Spanish1290—

Instruction Following Too close to call

Claude 3.5 Sonnet: 68.8 (#182), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
LMArena Instruction Following12971305
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Qwen3 32B leads

Claude 3.5 Sonnet: 39.9 (#167), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
LMArena Longer Query13111327
Fiction.LiveBench—74.2%

Writing & Preference Too close to call

Claude 3.5 Sonnet: 52.9 (#164), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetQwen3 32B
LMArena Text12981340
LMArena Creative Writing12921297
LMArena Multi-Turn13261331
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Qwen3 32B better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 37.7 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Qwen3 32B share?

18 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper