Model comparison

Claude 3.5 Haiku vs Qwen3 8B

Qwen3 8B is the stronger model overall, scoring 33.7 to 29.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 2 categories and Qwen3 8B in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3 8B leads 34.9 to 14.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 56.1% for Qwen3 8B.
  • Qwen3 8B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Qwen3 8B specifications
Claude 3.5 HaikuQwen3 8B
ProviderAnthropicAlibaba (Qwen)
Noometry Index29.233.7
Released2024-10-222025-04
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked4911

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 8B leads

Claude 3.5 Haiku: 32.9 (#265), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
SciCode27.4%22.6%
Aider Polyglot28%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
LMArena Coding1286—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Qwen3 8B leads

Claude 3.5 Haiku: 28.0 (#95), Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
Berkeley Function Calling Leaderboard—42.6%
BALROG19.3%—

Reasoning Claude 3.5 Haiku leads

Claude 3.5 Haiku: 17.7 (#290), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
CritPt0%0%
DTBench56.7%59.7%
Epoch Capabilities Index127.15136.17
Chess Puzzles—5%
LiveBench Reasoning28.1%—
LMArena Hard Prompts1251—
LiveBench Data Analysis48.5%—
LMCA—8.8%
LiveBench43.5%—

Math Qwen3 8B leads

Claude 3.5 Haiku: 14.7 (#300), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
OTIS Mock AIME 2024-20254.3%56.1%
Omni-MATH22.4%—
LiveBench Math35.5%—
LMArena Math1244—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen3 8B leads

Claude 3.5 Haiku: 18.7 (#281), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
GPQA Diamond38.1%56.8%
MMLU-Pro60.5%—
Confabulations36.7%—
Vectara Hallucination Rate—4.8%
GPQA (HELM)36.3%—
LMArena Expert1208—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Qwen3 8B: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
LMArena Vision1092—
GeoBench34%—

Multilingual Not comparable

Claude 3.5 Haiku: 40.0 (#218), Qwen3 8B: —

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
LMArena Non-English1238—
LMArena Chinese1229—
LMArena French1264—
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Russian1253—
LMArena Spanish1261—

Instruction Following Not comparable

Claude 3.5 Haiku: 62.9 (#234), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
LiveBench Instruction Following61.9%—
IFEval79.2%—
LMArena Instruction Following1241—

Long Context Too close to call

Claude 3.5 Haiku: 38.3 (#200), Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
Fiction.LiveBench—62.1%
LMArena Longer Query1261—

Writing & Preference Not comparable

Claude 3.5 Haiku: 42.7 (#234), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuQwen3 8B
LMArena Text1255—
LMArena Creative Writing1233—
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LMArena Multi-Turn1265—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Qwen3 8B?

Qwen3 8B is the stronger model overall, scoring 33.7 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Qwen3 8B better for coding?

Qwen3 8B scores higher on coding benchmarks: 34.0 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Qwen3 8B share?

6 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper