Model comparison

Claude 2.1 vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Qwen Max in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Max leads 30.3 to 15.4.
  • The biggest single-benchmark swing is GPQA Diamond: 33% for Claude 2.1 and 56.1% for Qwen Max.

Side by side

Claude 2.1 and Qwen Max specifications
Claude 2.1Qwen Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index25.234.7
Released2023-11-212024-04-03
WeightsProprietaryProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked723

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Claude 2.1: 26.2 (#327), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkClaude 2.1Qwen Max
Aider Polyglot—21.8%
WeirdML7.1%—
LMArena Coding—1288

Reasoning Qwen Max leads

Claude 2.1: 21.4 (#221), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkClaude 2.1Qwen Max
LMArena Hard Prompts—1269
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Qwen Max leads

Claude 2.1: 10.2 (#315), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkClaude 2.1Qwen Max
OTIS Mock AIME 2024-20251.9%16.1%
LMArena Math—1275
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Qwen Max leads

Claude 2.1: 15.4 (#292), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkClaude 2.1Qwen Max
GPQA Diamond33%56.1%
LMArena Expert—1248
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkClaude 2.1Qwen Max
LMArena Non-English—1263
LMArena Chinese—1254
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Russian—1274
LMArena Spanish—1290

Instruction Following Not comparable

Claude 2.1: —, Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkClaude 2.1Qwen Max
LMArena Instruction Following—1262

Long Context Not comparable

Claude 2.1: —, Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkClaude 2.1Qwen Max
Fiction.LiveBench—66.7%
LMArena Longer Query—1288

Writing & Preference Not comparable

Claude 2.1: —, Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkClaude 2.1Qwen Max
LMArena Text—1282
LMArena Creative Writing—1248
LMArena Multi-Turn—1277

Frequently asked questions

Is Claude 2.1 better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 25.2 on the Noometry Index.

Is Claude 2.1 or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Qwen Max share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper