Model comparison

Claude 3.7 Sonnet vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 39.5 on the Noometry Index.

Last verified . 22 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 1 category and Qwen3.6 Plus in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Plus leads 56.1 to 39.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Claude 3.7 Sonnet and 93.3% for Qwen3.6 Plus.

Side by side

Claude 3.7 Sonnet and Qwen3.6 Plus specifications
Claude 3.7 SonnetQwen3.6 Plus
ProviderAnthropicAlibaba (Qwen)
Noometry Index39.547.5
Released2025-02-242026-03-31
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.50
Output $ / M tokens—$3
Results tracked5837

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.7 Sonnet: 40.6 (#136), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
SWE-bench Verified61%57.9%
LMArena Coding13611467
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
LMArena WebDev—1461
SciCode—40.7%
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—
ALE-Bench—670.15

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—
Vending-Bench 2—5,115

Reasoning Qwen3.6 Plus leads

Claude 3.7 Sonnet: 18.6 (#277), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
LMArena Hard Prompts13331449
Epoch Capabilities Index141.16147.65
ARC-AGI-20.9%—
SimpleBench46.4%—
NYT Connections (extended)—60.3%
ARC-AGI-128.6%—
CritPt—2.9%
Chess Puzzles—17%
EnigmaEval4.2%—
Thematic Generalization—59.5%
LiveBench Reasoning87.8%—
Mystery Game Puzzles—12%
DTBench—81.9%
LiveBench Data Analysis74%—
LMCA—33.1%
ForecastBench61.8—
LiveBench76.1%—

Math Qwen3.6 Plus leads

Claude 3.7 Sonnet: 37.5 (#153), Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
OTIS Mock AIME 2024-202557.8%93.3%
LMArena Math13371450
FrontierMath (Feb 2025 set)4.1%26.2%
FrontierMath (Tiers 1-3)—38.2%
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath Tier 4 (v1)—8.3%

Knowledge Qwen3.6 Plus leads

Claude 3.7 Sonnet: 39.8 (#130), Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
GPQA Diamond79.7%88.4%
LMArena Expert13211454
Humanity's Last Exam8%—
SimpleQA Verified—44.1%
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Qwen3.6 Plus: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Qwen3.6 Plus leads

Claude 3.7 Sonnet: 44.1 (#179), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
LMArena Non-English12961424
LMArena Chinese12991477
LMArena French13031455
LMArena German13011452
LMArena Japanese12671389
LMArena Korean12491379
LMArena Russian13111434
LMArena Spanish12981432

Instruction Following Qwen3.6 Plus leads

Claude 3.7 Sonnet: 72.9 (#125), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
LMArena Instruction Following13521425
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
LMArena Longer Query13731439
Fiction.LiveBench83.3%—
CL-bench—20.3%

Writing & Preference Qwen3.6 Plus leads

Claude 3.7 Sonnet: 54.4 (#150), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetQwen3.6 Plus
LMArena Text13141437
LMArena Creative Writing13321404
LMArena Multi-Turn13391438
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or Qwen3.6 Plus better for coding?

They score almost the same on coding (40.6 vs 40.8); test both on your own repository before choosing.

How many benchmarks do Claude 3.7 Sonnet and Qwen3.6 Plus share?

22 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper