Model comparison

Qwen2.5 72B Instruct vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Qwen2.5 72B Instruct scores higher in 0 categories and Qwen3.5 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.5 Max Preview leads 40.1 to 19.3.
  • Qwen2.5 72B Instruct has downloadable open weights; the other is API-only.

Side by side

Qwen2.5 72B Instruct and Qwen3.5 Max Preview specifications
Qwen2.5 72B InstructQwen3.5 Max Preview
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index31.945.3
Released2024-09—
WeightsOpenProprietary
Context window131K—
Max output8K—
Input $ / M tokens$1.40—
Output $ / M tokens$5.60—
Results tracked4317

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 33.2 (#260), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Coding12921487
WeirdML16%—
BigCodeBench Instruct45.8%—
BigCodeBench Complete55.9%—

Agentic & Tool Use Not comparable

Qwen2.5 72B Instruct: 22.1 (#133), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
TheAgentCompany5.7%—
BALROG16.2%—
METR Time Horizons35.8%—

Reasoning Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 22.3 (#199), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Hard Prompts12711483
DTBench62.9%—
LMCA13.4%—
BIG-Bench Hard79.8%—
Epoch Capabilities Index129—
ForecastBench57.5—
HellaSwag84.8%—
PIQA82.6%—
WinoGrande82.3%—

Math Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 19.3 (#287), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Math12831474
OTIS Mock AIME 2024-20258.1%—
Omni-MATH33%—
MATH Level 563.2%—

Knowledge Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 27.0 (#253), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Expert12451489
GPQA Diamond49.1%—
MMLU-Pro63.1%—
Confabulations19.1%—
GPQA (HELM)42.6%—
ARC (AI2) Challenge94.5%—
MMLU85.3%—
TriviaQA71.9%—

Multilingual Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 41.0 (#213), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Non-English12521465
LMArena Chinese12721534
LMArena French12801484
LMArena German12341487
LMArena Japanese11801495
LMArena Korean11881438
LMArena Russian12641471
LMArena Spanish12561470

Instruction Following Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 65.5 (#221), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Instruction Following12541467
IFEval80.6%—

Long Context Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 38.9 (#188), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Longer Query12821476

Writing & Preference Qwen3.5 Max Preview leads

Qwen2.5 72B Instruct: 46.7 (#215), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkQwen2.5 72B InstructQwen3.5 Max Preview
LMArena Text12691470
LMArena Creative Writing12211464
LMArena Multi-Turn12721478
WildBench80.2%—

Frequently asked questions

Is Qwen2.5 72B Instruct better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 31.9 on the Noometry Index.

Is Qwen2.5 72B Instruct or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 33.2 in the Noometry coding category.

How many benchmarks do Qwen2.5 72B Instruct and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Qwen2.5 72B Instruct has 43 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper