Model comparison

Claude 3.5 Haiku vs Qwen2.5-Max

Qwen2.5-Max is the stronger model overall, scoring 40.7 to 29.2 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Qwen2.5-Max Alibaba (Qwen)

40.7

Rank #146 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Qwen2.5-Max in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen2.5-Max leads 36.9 to 14.7.
  • The biggest single-benchmark swing is LiveBench Reasoning: 28.1% for Claude 3.5 Haiku and 51.4% for Qwen2.5-Max.

Side by side

Claude 3.5 Haiku and Qwen2.5-Max specifications
Claude 3.5 HaikuQwen2.5-Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index29.240.7
Released2024-10-222025-01-25
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5-Max leads

Claude 3.5 Haiku: 32.9 (#265), Qwen2.5-Max: 41.8 (#117)

Coding benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LiveBench Coding51.4%64.4%
LMArena Coding12861359
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Qwen2.5-Max: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
BALROG19.3%—

Reasoning Qwen2.5-Max leads

Claude 3.5 Haiku: 17.7 (#290), Qwen2.5-Max: 25.6 (#147)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LiveBench Reasoning28.1%51.4%
LMArena Hard Prompts12511360
LiveBench Data Analysis48.5%67.9%
Epoch Capabilities Index127.15132.53
LiveBench43.5%62.3%
CritPt0%—
DTBench56.7%—

Math Qwen2.5-Max leads

Claude 3.5 Haiku: 14.7 (#300), Qwen2.5-Max: 36.9 (#162)

Math benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LiveBench Math35.5%58.4%
LMArena Math12441369
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen2.5-Max leads

Claude 3.5 Haiku: 18.7 (#281), Qwen2.5-Max: 35.3 (#186)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
Confabulations36.7%21.8%
LMArena Expert12081337
GPQA Diamond38.1%—
MMLU-Pro60.5%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Qwen2.5-Max: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LMArena Vision1092—
GeoBench34%—

Multilingual Qwen2.5-Max leads

Claude 3.5 Haiku: 40.0 (#218), Qwen2.5-Max: 48.1 (#146)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LMArena Non-English12381352
LMArena Chinese12291382
LMArena French12641396
LMArena German12371350
LMArena Japanese11751300
LMArena Korean11731304
LMArena Russian12531353
LMArena Spanish12611377

Instruction Following Qwen2.5-Max leads

Claude 3.5 Haiku: 62.9 (#234), Qwen2.5-Max: 71.3 (#152)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LiveBench Instruction Following61.9%75.3%
LMArena Instruction Following12411335
IFEval79.2%—

Long Context Qwen2.5-Max leads

Claude 3.5 Haiku: 38.3 (#200), Qwen2.5-Max: 41.4 (#142)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LMArena Longer Query12611358

Writing & Preference Qwen2.5-Max leads

Claude 3.5 Haiku: 42.7 (#234), Qwen2.5-Max: 55.4 (#146)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuQwen2.5-Max
LMArena Text12551367
LMArena Creative Writing12331339
Short-Story Creative Writing73.5%72.9%
LMArena Multi-Turn12651364
LiveBench Language35.4%56.3%
EQ-Bench Creative Writing1146—
WildBench76%—

Frequently asked questions

Is Claude 3.5 Haiku better than Qwen2.5-Max?

Qwen2.5-Max is the stronger model overall, scoring 40.7 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Qwen2.5-Max better for coding?

Qwen2.5-Max scores higher on coding benchmarks: 41.8 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Qwen2.5-Max share?

27 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Qwen2.5-Max has 27.

Related comparisons

Go deeper