Model comparison

Qwen Plus vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 37.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Qwen Plus scores higher in 6 categories and Qwen2.5 Plus 1127 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen2.5 Plus 1127 leads 36.1 to 23.3.

Side by side

Qwen Plus and Qwen2.5 Plus 1127 specifications
Qwen PlusQwen2.5 Plus 1127
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index37.138.8
Released2024-01-25—
WeightsProprietaryProprietary
Context window1M—
Max output33K—
Input $ / M tokens$0.40—
Output $ / M tokens$1.20—
Results tracked2014

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Qwen Plus: 38.9 (#167), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Coding13281314

Reasoning Qwen Plus leads

Qwen Plus: 28.4 (#107), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Hard Prompts13171299
Kagi LLM Benchmark63.3%—
DTBench81.1%—
LMCA24%—

Math Qwen2.5 Plus 1127 leads

Qwen Plus: 23.3 (#271), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Math13261298
OTIS Mock AIME 2024-202517.8%—
MATH Level 565.3%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Qwen2.5 Plus 1127 leads

Qwen Plus: 27.4 (#251), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Expert13281289
GPQA Diamond48.1%—

Multilingual Qwen Plus leads

Qwen Plus: 45.1 (#175), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Non-English13101265
LMArena Chinese13471314
LMArena Japanese12511207
LMArena Russian13231271
LMArena German—1231

Instruction Following Qwen Plus leads

Qwen Plus: 68.8 (#181), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Instruction Following13031275

Long Context Qwen Plus leads

Qwen Plus: 40.3 (#158), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Longer Query13241292

Writing & Preference Qwen Plus leads

Qwen Plus: 52.2 (#176), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkQwen PlusQwen2.5 Plus 1127
LMArena Text13261299
LMArena Creative Writing12931262
LMArena Multi-Turn13361299

Frequently asked questions

Is Qwen Plus better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 37.1 on the Noometry Index.

Is Qwen Plus or Qwen2.5 Plus 1127 better for coding?

They score almost the same on coding (38.9 vs 38.5); test both on your own repository before choosing.

How many benchmarks do Qwen Plus and Qwen2.5 Plus 1127 share?

13 benchmarks have published results for both models. Qwen Plus has 20 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper