Model comparison

Qwen Max vs Qwen2.5 72B Instruct

Qwen Max is the stronger model overall, scoring 34.7 to 31.9 on the Noometry Index.

Last verified . 20 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Qwen Max scores higher in 7 categories and Qwen2.5 72B Instruct in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Max leads 30.3 to 27.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 16.1% for Qwen Max and 8.1% for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct is cheaper at $1.40 / $5.60 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Qwen2.5 72B Instruct accepts more context: 131K tokens versus 33K.
  • Qwen2.5 72B Instruct has downloadable open weights; the other is API-only.

Side by side

Qwen Max and Qwen2.5 72B Instruct specifications
Qwen MaxQwen2.5 72B Instruct
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.731.9
Released2024-04-032024-09
WeightsProprietaryOpen
Context window33K131K
Max output8K8K
Input $ / M tokens$1.60$1.40
Output $ / M tokens$6.40$5.60
Results tracked2343

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 72B Instruct leads

Qwen Max: 30.7 (#292), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
LMArena Coding12881292
Aider Polyglot21.8%—
WeirdML—16%
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%

Agentic & Tool Use Not comparable

Qwen Max: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
LMArena Hard Prompts12691271
DTBench—62.9%
LMCA—13.4%
BIG-Bench Hard—79.8%
Epoch Capabilities Index—129
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Qwen Max leads

Qwen Max: 22.3 (#276), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
OTIS Mock AIME 2024-202516.1%8.1%
LMArena Math12751283
MATH Level 567.2%63.2%
Omni-MATH—33%
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen Max leads

Qwen Max: 30.3 (#228), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
GPQA Diamond56.1%49.1%
LMArena Expert12481245
MMLU-Pro—63.1%
Confabulations—19.1%
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multilingual Too close to call

Qwen Max: 41.8 (#202), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
LMArena Non-English12631252
LMArena Chinese12541272
LMArena French13301280
LMArena German12541234
LMArena Japanese12051180
LMArena Korean11421188
LMArena Russian12741264
LMArena Spanish12901256

Instruction Following Qwen Max leads

Qwen Max: 66.5 (#208), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
LMArena Instruction Following12621254
IFEval—80.6%

Long Context Too close to call

Qwen Max: 39.4 (#180), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
LMArena Longer Query12881282
Fiction.LiveBench66.7%—

Writing & Preference Qwen Max leads

Qwen Max: 47.8 (#205), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkQwen MaxQwen2.5 72B Instruct
LMArena Text12821269
LMArena Creative Writing12481221
LMArena Multi-Turn12771272
WildBench—80.2%

Frequently asked questions

Is Qwen Max better than Qwen2.5 72B Instruct?

Qwen Max is the stronger model overall, scoring 34.7 to 31.9 on the Noometry Index.

Which is cheaper, Qwen Max or Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is cheaper. It lists at $1.40 per million input tokens and $5.60 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Qwen Max or Qwen2.5 72B Instruct better for coding?

Qwen2.5 72B Instruct scores higher on coding benchmarks: 33.2 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Qwen2.5 72B Instruct does, with 131K tokens against 33K.

How many benchmarks do Qwen Max and Qwen2.5 72B Instruct share?

20 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper