Model comparison

Claude 3.5 Sonnet vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 34.6 on the Noometry Index.

Last verified . 18 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 3 categories and Qwen Plus in 5 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen Plus leads 28.4 to 23.1.
  • The biggest single-benchmark swing is DTBench: 67.8% for Claude 3.5 Sonnet and 81.1% for Qwen Plus.

Side by side

Claude 3.5 Sonnet and Qwen Plus specifications
Claude 3.5 SonnetQwen Plus
ProviderAnthropicAlibaba (Qwen)
Noometry Index34.637.1
Released2024-06-202024-01-25
WeightsProprietaryProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked6020

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Coding13421328
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Qwen Plus: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Qwen Plus leads

Claude 3.5 Sonnet: 23.1 (#183), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Hard Prompts13051317
DTBench67.8%81.1%
SimpleBench41.4%—
Kagi LLM Benchmark—63.3%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LiveBench Data Analysis55%—
LMCA—24%
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Qwen Plus leads

Claude 3.5 Sonnet: 19.2 (#288), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
OTIS Mock AIME 2024-20258.5%17.8%
LMArena Math13071326
MATH Level 556.9%65.3%
FrontierMath (Feb 2025 set)2.1%1.7%
Omni-MATH27.6%—
LiveBench Math52.3%—
FrontierMath Tier 4 (v1)0%—

Knowledge Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 28.6 (#245), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
GPQA Diamond55.3%48.1%
LMArena Expert12651328
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Qwen Plus: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Qwen Plus leads

Claude 3.5 Sonnet: 43.2 (#185), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Non-English12831310
LMArena Chinese12721347
LMArena Japanese12341251
LMArena Russian13061323
LMArena French1305—
LMArena German1297—
LMArena Korean1200—
LMArena Spanish1290—

Instruction Following Too close to call

Claude 3.5 Sonnet: 68.8 (#182), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Instruction Following12971303
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Longer Query13111324

Writing & Preference Too close to call

Claude 3.5 Sonnet: 52.9 (#164), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetQwen Plus
LMArena Text12981326
LMArena Creative Writing12921293
LMArena Multi-Turn13261336
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Qwen Plus better for coding?

They score almost the same on coding (39.0 vs 38.9); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and Qwen Plus share?

18 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper