Model comparison

Qwen Plus vs Qwen2.5 72B Instruct

Qwen Plus is the stronger model overall, scoring 37.1 to 31.9 on the Noometry Index.

Last verified . 18 shared benchmarks.

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Qwen Plus scores higher in 8 categories and Qwen2.5 72B Instruct in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen Plus leads 28.4 to 22.3.
  • The biggest single-benchmark swing is DTBench: 81.1% for Qwen Plus and 62.9% for Qwen2.5 72B Instruct.
  • Qwen Plus is cheaper at $0.40 / $1.20 per million input/output tokens, against $1.40 / $5.60 for Qwen2.5 72B Instruct.
  • Qwen Plus accepts more context: 1M tokens versus 131K.
  • Qwen2.5 72B Instruct has downloadable open weights; the other is API-only.

Side by side

Qwen Plus and Qwen2.5 72B Instruct specifications
Qwen PlusQwen2.5 72B Instruct
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index37.131.9
Released2024-01-252024-09
WeightsProprietaryOpen
Context window1M131K
Max output33K8K
Input $ / M tokens$0.40$1.40
Output $ / M tokens$1.20$5.60
Results tracked2043

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Qwen Plus: 38.9 (#167), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
LMArena Coding13281292
WeirdML—16%
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%

Agentic & Tool Use Not comparable

Qwen Plus: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen Plus leads

Qwen Plus: 28.4 (#107), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
LMArena Hard Prompts13171271
DTBench81.1%62.9%
LMCA24%13.4%
Kagi LLM Benchmark63.3%—
BIG-Bench Hard—79.8%
Epoch Capabilities Index—129
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Qwen Plus leads

Qwen Plus: 23.3 (#271), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
OTIS Mock AIME 2024-202517.8%8.1%
LMArena Math13261283
MATH Level 565.3%63.2%
Omni-MATH—33%
FrontierMath (Feb 2025 set)1.7%—

Knowledge Too close to call

Qwen Plus: 27.4 (#251), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
GPQA Diamond48.1%49.1%
LMArena Expert13281245
MMLU-Pro—63.1%
Confabulations—19.1%
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multilingual Qwen Plus leads

Qwen Plus: 45.1 (#175), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
LMArena Non-English13101252
LMArena Chinese13471272
LMArena Japanese12511180
LMArena Russian13231264
LMArena French—1280
LMArena German—1234
LMArena Korean—1188
LMArena Spanish—1256

Instruction Following Qwen Plus leads

Qwen Plus: 68.8 (#181), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
LMArena Instruction Following13031254
IFEval—80.6%

Long Context Qwen Plus leads

Qwen Plus: 40.3 (#158), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
LMArena Longer Query13241282

Writing & Preference Qwen Plus leads

Qwen Plus: 52.2 (#176), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkQwen PlusQwen2.5 72B Instruct
LMArena Text13261269
LMArena Creative Writing12931221
LMArena Multi-Turn13361272
WildBench—80.2%

Frequently asked questions

Is Qwen Plus better than Qwen2.5 72B Instruct?

Qwen Plus is the stronger model overall, scoring 37.1 to 31.9 on the Noometry Index.

Which is cheaper, Qwen Plus or Qwen2.5 72B Instruct?

Qwen Plus is cheaper. It lists at $0.40 per million input tokens and $1.20 per million output tokens; Qwen2.5 72B Instruct lists at $1.40 and $5.60.

Is Qwen Plus or Qwen2.5 72B Instruct better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 33.2 in the Noometry coding category.

Which has the bigger context window?

Qwen Plus does, with 1M tokens against 131K.

How many benchmarks do Qwen Plus and Qwen2.5 72B Instruct share?

18 benchmarks have published results for both models. Qwen Plus has 20 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper