Model comparison

Claude 3.5 Haiku vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 29.2 on the Noometry Index.

Last verified . 20 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Qwen3 32B in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3 32B leads 39.7 to 14.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 66.9% for Qwen3 32B.
  • Qwen3 32B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Qwen3 32B specifications
Claude 3.5 HaikuQwen3 32B
ProviderAnthropicAlibaba (Qwen)
Noometry Index29.239.2
Released2024-10-222025-04
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked4926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 32B leads

Claude 3.5 Haiku: 32.9 (#265), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
Aider Polyglot28%40%
SciCode27.4%35.4%
LMArena Coding12861358
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Qwen3 32B leads

Claude 3.5 Haiku: 28.0 (#95), Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
Berkeley Function Calling Leaderboard—48.7%
BALROG19.3%—

Reasoning Qwen3 32B leads

Claude 3.5 Haiku: 17.7 (#290), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
CritPt0%0.3%
LMArena Hard Prompts12511334
DTBench56.7%67.5%
Epoch Capabilities Index127.15138.51
Kagi LLM Benchmark—54.9%
Chess Puzzles—5%
LiveBench Reasoning28.1%—
LiveBench Data Analysis48.5%—
LMCA—17.3%
LiveBench43.5%—

Math Qwen3 32B leads

Claude 3.5 Haiku: 14.7 (#300), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
OTIS Mock AIME 2024-20254.3%66.9%
LMArena Math12441399
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen3 32B leads

Claude 3.5 Haiku: 18.7 (#281), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
GPQA Diamond38.1%65.7%
LMArena Expert12081362
MMLU-Pro60.5%—
Confabulations36.7%—
Vectara Hallucination Rate—5.9%
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Qwen3 32B: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
LMArena Vision1092—
GeoBench34%—

Multilingual Qwen3 32B leads

Claude 3.5 Haiku: 40.0 (#218), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
LMArena Non-English12381317
LMArena Chinese12291357
LMArena German12371341
LMArena Russian12531311
LMArena French1264—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Spanish1261—

Instruction Following Qwen3 32B leads

Claude 3.5 Haiku: 62.9 (#234), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
LMArena Instruction Following12411305
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Qwen3 32B leads

Claude 3.5 Haiku: 38.3 (#200), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
LMArena Longer Query12611327
Fiction.LiveBench—74.2%

Writing & Preference Qwen3 32B leads

Claude 3.5 Haiku: 42.7 (#234), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuQwen3 32B
LMArena Text12551340
LMArena Creative Writing12331297
LMArena Multi-Turn12651331
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Qwen3 32B better for coding?

Qwen3 32B scores higher on coding benchmarks: 37.7 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Qwen3 32B share?

20 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper