Model comparison

Claude Haiku 4.5 vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 39.5 on the Noometry Index.

Last verified . 28 shared benchmarks.

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Claude Haiku 4.5 scores higher in 3 categories and Qwen3 Max in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 37.7.
  • The biggest single-benchmark swing is SimpleQA Verified: 13.2% for Claude Haiku 4.5 and 48.7% for Qwen3 Max.
  • Claude Haiku 4.5 is cheaper at $1 / $5 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Qwen3 Max accepts more context: 262K tokens versus 200K.

Side by side

Claude Haiku 4.5 and Qwen3 Max specifications
Claude Haiku 4.5Qwen3 Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index39.543.7
Released2025-10-152025-09-23
WeightsProprietaryProprietary
Context window200K262K
Max output64K66K
Input $ / M tokens$1$1.20
Output $ / M tokens$5$6
Results tracked5333

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.0 (#78), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
LMArena Coding14531456
ALE-Bench653.48370.45
SWE-bench Verified (bash only)66.6%—
LMArena WebDev1330—
SWE-bench Multilingual64.7%—
SciCode43.3%—
WeirdML45.4%—

Agentic & Tool Use Not comparable

Claude Haiku 4.5: 33.6 (#52), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
Vending-Bench 2458.8971.56
Terminal-Bench35.5%—
Berkeley Function Calling Leaderboard68.7%—
DeepResearch Bench45.5%—
BALROG31.2%—
ExploitBench13.7%—

Reasoning Qwen3 Max leads

Claude Haiku 4.5: 15.1 (#320), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
NYT Connections (extended)14.3%30.1%
Chess Puzzles8%4%
LMArena Hard Prompts14201448
DTBench73.6%82.1%
LMCA30.9%28.3%
Epoch Capabilities Index142.41142.38
ARC-AGI-24%—
Kagi LLM Benchmark—72.5%
ARC-AGI-147.7%—
CritPt0%—
Mystery Game Puzzles—5%
ForecastBench61.4—

Math Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.9 (#78), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
OTIS Mock AIME 2024-202566.7%73.3%
LMArena Math13961446
MATH Level 596.4%97.1%
FrontierMath (Tiers 1-3)—18.9%
Omni-MATH56.1%—
FrontierMath (Feb 2025 set)5.9%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Qwen3 Max leads

Claude Haiku 4.5: 37.7 (#153), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
GPQA Diamond71.2%72.6%
SimpleQA Verified13.2%48.7%
LMArena Expert14421455
MMLU-Pro77.7%—
Vectara Hallucination Rate9.8%—
GPQA (HELM)60.5%—

Multimodal Not comparable

Claude Haiku 4.5: 26.8 (#118), Qwen3 Max: —

Multimodal benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
Blueprint-Bench 20%—
LMArena Document1420—

Multilingual Qwen3 Max leads

Claude Haiku 4.5: 49.9 (#129), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
LMArena Non-English13771429
LMArena Chinese14171478
LMArena French14081449
LMArena German13751463
LMArena Japanese13391397
LMArena Korean13471399
LMArena Russian13811428
LMArena Spanish14201462

Instruction Following Qwen3 Max leads

Claude Haiku 4.5: 71.4 (#149), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
LMArena Instruction Following14141419
IFEval80.1%—

Long Context Claude Haiku 4.5 leads

Claude Haiku 4.5: 43.6 (#92), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
LMArena Longer Query14271438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Claude Haiku 4.5: 57.9 (#123), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkClaude Haiku 4.5Qwen3 Max
LMArena Text13961439
LMArena Creative Writing13721402
LMArena Multi-Turn14091446
WildBench83.9%—
EQ-Bench 41064—

Frequently asked questions

Is Claude Haiku 4.5 better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 39.5 on the Noometry Index.

Which is cheaper, Claude Haiku 4.5 or Qwen3 Max?

Claude Haiku 4.5 is cheaper. It lists at $1 per million input tokens and $5 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Claude Haiku 4.5 or Qwen3 Max better for coding?

Claude Haiku 4.5 scores higher on coding benchmarks: 44.0 versus 43.0 in the Noometry coding category.

Which has the bigger context window?

Qwen3 Max does, with 262K tokens against 200K.

How many benchmarks do Claude Haiku 4.5 and Qwen3 Max share?

28 benchmarks have published results for both models. Claude Haiku 4.5 has 53 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper