Model comparison

Claude 3.7 Sonnet vs Qwen Plus

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 37.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 6 categories and Qwen Plus in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude 3.7 Sonnet leads 37.5 to 23.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Claude 3.7 Sonnet and 17.8% for Qwen Plus.

Side by side

Claude 3.7 Sonnet and Qwen Plus specifications
Claude 3.7 SonnetQwen Plus
ProviderAnthropicAlibaba (Qwen)
Noometry Index39.537.1
Released2025-02-242024-01-25
WeightsProprietaryProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked5820

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Coding13611328
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Qwen Plus: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Qwen Plus leads

Claude 3.7 Sonnet: 18.6 (#277), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Hard Prompts13331317
ARC-AGI-20.9%—
SimpleBench46.4%—
Kagi LLM Benchmark—63.3%
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
DTBench—81.1%
LiveBench Data Analysis74%—
LMCA—24%
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
OTIS Mock AIME 2024-202557.8%17.8%
LMArena Math13371326
MATH Level 591.2%65.3%
FrontierMath (Feb 2025 set)4.1%1.7%
Omni-MATH33%—
LiveBench Math79%—

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
GPQA Diamond79.7%48.1%
LMArena Expert13211328
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Qwen Plus: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Too close to call

Claude 3.7 Sonnet: 44.1 (#179), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Non-English12961310
LMArena Chinese12991347
LMArena Japanese12671251
LMArena Russian13111323
LMArena French1303—
LMArena German1301—
LMArena Korean1249—
LMArena Spanish1298—

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Instruction Following13521303
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Longer Query13731324
Fiction.LiveBench83.3%—

Writing & Preference Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 54.4 (#150), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetQwen Plus
LMArena Text13141326
LMArena Creative Writing13321293
LMArena Multi-Turn13391336
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Qwen Plus?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 37.1 on the Noometry Index.

Is Claude 3.7 Sonnet or Qwen Plus better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 38.9 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Qwen Plus share?

17 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper