Model comparison

Claude 3.5 Haiku vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 29.2 on the Noometry Index.

Last verified . 19 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Qwen3.6 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.6 Max Preview leads 54.1 to 14.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 91.1% for Qwen3.6 Max Preview.

Side by side

Claude 3.5 Haiku and Qwen3.6 Max Preview specifications
Claude 3.5 HaikuQwen3.6 Max Preview
ProviderAnthropicAlibaba (Qwen)
Noometry Index29.251.5
Released2024-10-222026-04-20
WeightsProprietaryProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.30
Output $ / M tokens—$7.80
Results tracked4929

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 32.9 (#265), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Coding12861471
SWE-bench Verified—76.7%
Aider Polyglot28%—
LMArena WebDev—1482
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
BALROG19.3%—
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 17.7 (#290), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Hard Prompts12511457
DTBench56.7%87.2%
Epoch Capabilities Index127.15149.24
SimpleBench—63%
NYT Connections (extended)—74.1%
CritPt0%—
Chess Puzzles—20%
LiveBench Reasoning28.1%—
Mystery Game Puzzles—19%
LiveBench Data Analysis48.5%—
LMCA—42.5%
LiveBench43.5%—

Math Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 14.7 (#300), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
OTIS Mock AIME 2024-20254.3%91.1%
LMArena Math12441465
FrontierMath (Feb 2025 set)0.3%23.1%
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath Tier 4 (v1)—4.2%

Knowledge Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 18.7 (#281), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
GPQA Diamond38.1%87.4%
LMArena Expert12081478
SimpleQA Verified—52%
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Vision1092—
GeoBench34%—

Multilingual Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 40.0 (#218), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Non-English12381437
LMArena Chinese12291487
LMArena French12641449
LMArena Russian12531445
LMArena Spanish12611454
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—

Instruction Following Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 62.9 (#234), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Instruction Following12411438
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 38.3 (#200), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Longer Query12611457

Writing & Preference Qwen3.6 Max Preview leads

Claude 3.5 Haiku: 42.7 (#234), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuQwen3.6 Max Preview
LMArena Text12551447
LMArena Creative Writing12331435
LMArena Multi-Turn12651456
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Qwen3.6 Max Preview share?

19 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper