Model comparison

Qwen Max vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 34.7 on the Noometry Index.

Last verified . 14 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Qwen Max scores higher in 1 category and Qwen2.5 Plus 1127 in 7 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen2.5 Plus 1127 leads 36.1 to 22.3.

Side by side

Qwen Max and Qwen2.5 Plus 1127 specifications
Qwen MaxQwen2.5 Plus 1127
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.738.8
Released2024-04-03—
WeightsProprietaryProprietary
Context window33K—
Max output8K—
Input $ / M tokens$1.60—
Output $ / M tokens$6.40—
Results tracked2314

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Qwen Max: 30.7 (#292), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Coding12881314
Aider Polyglot21.8%—

Reasoning Too close to call

Qwen Max: 25.1 (#151), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Hard Prompts12691299

Math Qwen2.5 Plus 1127 leads

Qwen Max: 22.3 (#276), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Math12751298
OTIS Mock AIME 2024-202516.1%—
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen2.5 Plus 1127 leads

Qwen Max: 30.3 (#228), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Expert12481289
GPQA Diamond56.1%—

Multilingual Too close to call

Qwen Max: 41.8 (#202), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Non-English12631265
LMArena Chinese12541314
LMArena German12541231
LMArena Japanese12051207
LMArena Russian12741271
LMArena French1330—
LMArena Korean1142—
LMArena Spanish1290—

Instruction Following Too close to call

Qwen Max: 66.5 (#208), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Instruction Following12621275

Long Context Too close to call

Qwen Max: 39.4 (#180), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Longer Query12881292
Fiction.LiveBench66.7%—

Writing & Preference Qwen2.5 Plus 1127 leads

Qwen Max: 47.8 (#205), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkQwen MaxQwen2.5 Plus 1127
LMArena Text12821299
LMArena Creative Writing12481262
LMArena Multi-Turn12771299

Frequently asked questions

Is Qwen Max better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 34.7 on the Noometry Index.

Is Qwen Max or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 30.7 in the Noometry coding category.

How many benchmarks do Qwen Max and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper