Model comparison

Claude 3.5 Haiku vs Qwen3.8 27B

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 29.2 on the Noometry Index.

Last verified . 23 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Qwen3.8 27B in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.8 27B leads 41.0 to 17.7.
  • The biggest single-benchmark swing is DTBench: 56.7% for Claude 3.5 Haiku and 88% for Qwen3.8 27B.
  • Qwen3.8 27B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Qwen3.8 27B specifications
Claude 3.5 HaikuQwen3.8 27B
ProviderAnthropicAlibaba (Qwen)
Noometry Index29.246.0
Released2024-10-222026-08-14
WeightsProprietaryOpen
Context window—262K
Max output—33K
Input $ / M tokens—$0.99
Output $ / M tokens—$1.49
Results tracked4931

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

Claude 3.5 Haiku: 32.9 (#265), Qwen3.8 27B: 50.5 (#44)

Coding benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
SciCode27.4%46.6%
LMArena Coding12861482
Aider Polyglot28%—
LMArena WebDev—1593
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Qwen3.8 27B leads

Claude 3.5 Haiku: 28.0 (#95), Qwen3.8 27B: 32.9 (#57)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
APEX-Agents—47.5%
BALROG19.3%—

Reasoning Qwen3.8 27B leads

Claude 3.5 Haiku: 17.7 (#290), Qwen3.8 27B: 41.0 (#54)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
CritPt0%5.4%
LMArena Hard Prompts12511460
DTBench56.7%88%
Epoch Capabilities Index127.15149.38
ARC-AGI-2—42.4%
NYT Connections (extended)—54.5%
ARC-AGI-1—87.5%
LiveBench Reasoning28.1%—
LiveBench Data Analysis48.5%—
LMCA—41.4%
Surface Evolver Bench—45%
LiveBench43.5%—

Math Qwen3.8 27B leads

Claude 3.5 Haiku: 14.7 (#300), Qwen3.8 27B: 37.1 (#161)

Math benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Math12441456
OTIS Mock AIME 2024-20254.3%—
ProofBench—16%
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen3.8 27B leads

Claude 3.5 Haiku: 18.7 (#281), Qwen3.8 27B: 41.6 (#109)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Expert12081482
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Qwen3.8 27B leads

Claude 3.5 Haiku: 26.8 (#117), Qwen3.8 27B: 41.3 (#37)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Vision10921271
GeoBench34%—

Multilingual Qwen3.8 27B leads

Claude 3.5 Haiku: 40.0 (#218), Qwen3.8 27B: 53.7 (#60)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Non-English12381430
LMArena Chinese12291504
LMArena French12641465
LMArena German12371438
LMArena Japanese11751384
LMArena Korean11731393
LMArena Russian12531415
LMArena Spanish12611448

Instruction Following Qwen3.8 27B leads

Claude 3.5 Haiku: 62.9 (#234), Qwen3.8 27B: 75.8 (#53)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Instruction Following12411439
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Qwen3.8 27B leads

Claude 3.5 Haiku: 38.3 (#200), Qwen3.8 27B: 44.3 (#70)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Longer Query12611450

Writing & Preference Qwen3.8 27B leads

Claude 3.5 Haiku: 42.7 (#234), Qwen3.8 27B: 65.8 (#43)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuQwen3.8 27B
LMArena Text12551441
LMArena Creative Writing12331384
EQ-Bench Creative Writing11461671
LMArena Multi-Turn12651441
Short-Story Creative Writing73.5%—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Qwen3.8 27B?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Qwen3.8 27B better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Qwen3.8 27B share?

23 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Qwen3.8 27B has 31.

Related comparisons

Go deeper