Model comparison

Claude 3.7 Sonnet vs Qwen Max

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 34.7 on the Noometry Index.

Last verified . 23 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 7 categories and Qwen Max in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude 3.7 Sonnet leads 37.5 to 22.3.
  • The biggest single-benchmark swing is Aider Polyglot: 64.9% for Claude 3.7 Sonnet and 21.8% for Qwen Max.

Side by side

Claude 3.7 Sonnet and Qwen Max specifications
Claude 3.7 SonnetQwen Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index39.534.7
Released2025-02-242024-04-03
WeightsProprietaryProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked5823

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
Aider Polyglot64.9%21.8%
LMArena Coding13611288
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Qwen Max leads

Claude 3.7 Sonnet: 18.6 (#277), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
LMArena Hard Prompts13331269
ARC-AGI-20.9%—
SimpleBench46.4%—
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
OTIS Mock AIME 2024-202557.8%16.1%
LMArena Math13371275
MATH Level 591.2%67.2%
FrontierMath (Feb 2025 set)4.1%1%
Omni-MATH33%—
LiveBench Math79%—

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
GPQA Diamond79.7%56.1%
LMArena Expert13211248
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Qwen Max: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 44.1 (#179), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
LMArena Non-English12961263
LMArena Chinese12991254
LMArena French13031330
LMArena German13011254
LMArena Japanese12671205
LMArena Korean12491142
LMArena Russian13111274
LMArena Spanish12981290

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
LMArena Instruction Following13521262
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
Fiction.LiveBench83.3%66.7%
LMArena Longer Query13731288

Writing & Preference Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 54.4 (#150), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetQwen Max
LMArena Text13141282
LMArena Creative Writing13321248
LMArena Multi-Turn13391277
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Qwen Max?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 34.7 on the Noometry Index.

Is Claude 3.7 Sonnet or Qwen Max better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 30.7 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Qwen Max share?

23 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper