Model comparison

Claude Sonnet 4.5 vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 44.1 on the Noometry Index.

Last verified . 35 shared benchmarks.

Claude Sonnet 4.5 Anthropic

44.1

Rank #81 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 35 benchmarks with published results for both. Claude Sonnet 4.5 scores higher in 4 categories and Qwen3.6 Plus in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.6 Plus leads 51.8 to 32.3.
  • The biggest single-benchmark swing is NYT Connections (extended): 37.3% for Claude Sonnet 4.5 and 60.3% for Qwen3.6 Plus.
  • Qwen3.6 Plus is cheaper at $0.50 / $3 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.5.
  • Qwen3.6 Plus accepts more context: 1M tokens versus 200K.

Side by side

Claude Sonnet 4.5 and Qwen3.6 Plus specifications
Claude Sonnet 4.5Qwen3.6 Plus
ProviderAnthropicAlibaba (Qwen)
Noometry Index44.147.5
Released2025-09-292026-03-31
WeightsProprietaryProprietary
Context window200K1M
Max output64K66K
Input $ / M tokens$3$0.50
Output $ / M tokens$15$3
Results tracked7337

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4.5 leads

Claude Sonnet 4.5: 47.3 (#61), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
SWE-bench Verified71.3%57.9%
LMArena WebDev13931461
SciCode44.7%40.7%
LMArena Coding14891467
ALE-Bench796.15670.15
SWE-bench Verified (bash only)71.4%—
SWE-bench Multilingual67%—
GSO14.7%—
WeirdML47.7%—
AlgoTune1.52—

Agentic & Tool Use Not comparable

Claude Sonnet 4.5: 38.3 (#32), Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
Vending-Bench 23,8395,115
Terminal-Bench46.5%—
Berkeley Function Calling Leaderboard73.2%—
GDPval42.5%—
Remote Labor Index2.1%—
τ²-bench Airline72%—
τ²-bench Banking25.3%—
τ²-bench Retail72.4%—
τ²-bench Telecom84.9%—
Cybench60%—
DeepResearch Bench52.6%—
OSWorld62.9%—
LMArena Search1159—
METR Time Horizons67.4%—

Reasoning Qwen3.6 Plus leads

Claude Sonnet 4.5: 26.9 (#125), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
NYT Connections (extended)37.3%60.3%
CritPt1.1%2.9%
Chess Puzzles12%17%
LMArena Hard Prompts14621449
Mystery Game Puzzles17%12%
DTBench83.2%81.9%
LMCA38.8%33.1%
Epoch Capabilities Index146.84147.65
ARC-AGI-213.6%—
SimpleBench54.3%—
Kagi LLM Benchmark57.9%—
ARC-AGI-163.7%—
EnigmaEval6%—
Thematic Generalization—59.5%
EBR-Bench2.4%—
ForecastBench61.9—

Math Qwen3.6 Plus leads

Claude Sonnet 4.5: 32.3 (#216), Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
FrontierMath (Tiers 1-3)23.9%38.2%
OTIS Mock AIME 2024-202577.8%93.3%
LMArena Math14491450
FrontierMath (Feb 2025 set)15.2%26.2%
FrontierMath Tier 4 (v1)4.2%8.3%
FrontierMath Tier 42.4%—
ProofBench19%—
Omni-MATH55.3%—
MATH Level 597.7%—

Knowledge Qwen3.6 Plus leads

Claude Sonnet 4.5: 48.4 (#76), Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
GPQA Diamond82.3%88.4%
SimpleQA Verified30.7%44.1%
LMArena Expert14821454
Humanity's Last Exam13.7%—
MMLU-Pro86.9%—
Vectara Hallucination Rate12%—
GPQA (HELM)68.6%—

Multimodal Not comparable

Claude Sonnet 4.5: 34.8 (#89), Qwen3.6 Plus: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
VPCT39.8%—
LMArena Document1450—

Multilingual Too close to call

Claude Sonnet 4.5: 53.4 (#69), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
LMArena Non-English14251424
LMArena Chinese14591477
LMArena French14581455
LMArena German14271452
LMArena Japanese13901389
LMArena Korean14031379
LMArena Russian14371434
LMArena Spanish14571432

Instruction Following Too close to call

Claude Sonnet 4.5: 75.0 (#78), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
LMArena Instruction Following14591425
IFEval85%—

Long Context Too close to call

Claude Sonnet 4.5: 45.2 (#46), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
LMArena Longer Query14761439
CL-bench—20.3%

Writing & Preference Claude Sonnet 4.5 leads

Claude Sonnet 4.5: 66.5 (#34), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.5Qwen3.6 Plus
LMArena Text14391437
LMArena Creative Writing14421404
LMArena Multi-Turn14651438
EQ-Bench Creative Writing1678—
WildBench85.4%—

Frequently asked questions

Is Claude Sonnet 4.5 better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 44.1 on the Noometry Index.

Which is cheaper, Claude Sonnet 4.5 or Qwen3.6 Plus?

Qwen3.6 Plus is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; Claude Sonnet 4.5 lists at $3 and $15.

Is Claude Sonnet 4.5 or Qwen3.6 Plus better for coding?

Claude Sonnet 4.5 scores higher on coding benchmarks: 47.3 versus 40.8 in the Noometry coding category.

Which has the bigger context window?

Qwen3.6 Plus does, with 1M tokens against 200K.

How many benchmarks do Claude Sonnet 4.5 and Qwen3.6 Plus share?

35 benchmarks have published results for both models. Claude Sonnet 4.5 has 73 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper