Model comparison

Claude 3.7 Sonnet vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 39.5 on the Noometry Index.

Last verified . 22 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 1 category and Qwen3 Max in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Qwen3 Max leads 53.7 to 44.1.
  • The biggest single-benchmark swing is Fiction.LiveBench: 83.3% for Claude 3.7 Sonnet and 66.7% for Qwen3 Max.

Side by side

Claude 3.7 Sonnet and Qwen3 Max specifications
Claude 3.7 SonnetQwen3 Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index39.543.7
Released2025-02-242025-09-23
WeightsProprietaryProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.20
Output $ / M tokens—$6
Results tracked5833

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 Max leads

Claude 3.7 Sonnet: 40.6 (#136), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
LMArena Coding13611456
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—
ALE-Bench—370.45

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—
Vending-Bench 2—71.56

Reasoning Qwen3 Max leads

Claude 3.7 Sonnet: 18.6 (#277), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
LMArena Hard Prompts13331448
Epoch Capabilities Index141.16142.38
ARC-AGI-20.9%—
SimpleBench46.4%—
Kagi LLM Benchmark—72.5%
NYT Connections (extended)—30.1%
ARC-AGI-128.6%—
Chess Puzzles—4%
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
Mystery Game Puzzles—5%
DTBench—82.1%
LiveBench Data Analysis74%—
LMCA—28.3%
ForecastBench61.8—
LiveBench76.1%—

Math Qwen3 Max leads

Claude 3.7 Sonnet: 37.5 (#153), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
OTIS Mock AIME 2024-202557.8%73.3%
LMArena Math13371446
MATH Level 591.2%97.1%
FrontierMath (Tiers 1-3)—18.9%
Omni-MATH33%—
LiveBench Math79%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Qwen3 Max leads

Claude 3.7 Sonnet: 39.8 (#130), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
GPQA Diamond79.7%72.6%
LMArena Expert13211455
Humanity's Last Exam8%—
SimpleQA Verified—48.7%
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Qwen3 Max: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Qwen3 Max leads

Claude 3.7 Sonnet: 44.1 (#179), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
LMArena Non-English12961429
LMArena Chinese12991478
LMArena French13031449
LMArena German13011463
LMArena Japanese12671397
LMArena Korean12491399
LMArena Russian13111428
LMArena Spanish12981462

Instruction Following Qwen3 Max leads

Claude 3.7 Sonnet: 72.9 (#125), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
LMArena Instruction Following13521419
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
Fiction.LiveBench83.3%66.7%
LMArena Longer Query13731438
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Claude 3.7 Sonnet: 54.4 (#150), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetQwen3 Max
LMArena Text13141439
LMArena Creative Writing13321402
LMArena Multi-Turn13391446
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or Qwen3 Max better for coding?

Qwen3 Max scores higher on coding benchmarks: 43.0 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Qwen3 Max share?

22 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper