Model comparison

Qwen Max vs Qwen2.5 32B Instruct

Qwen Max is the stronger model overall, scoring 34.7 to 30.1 on the Noometry Index. Qwen2.5 32B Instruct costs 2.3× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Last verified . 3 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen2.5 32B Instruct Alibaba (Qwen)

30.1

Rank #297 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Qwen Max scores higher in 3 categories and Qwen2.5 32B Instruct in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Qwen2.5 32B Instruct leads 38.7 to 30.7.
  • The biggest single-benchmark swing is MATH Level 5: 67.2% for Qwen Max and 56.1% for Qwen2.5 32B Instruct.
  • Qwen2.5 32B Instruct is cheaper at $0.70 / $2.80 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Qwen2.5 32B Instruct accepts more context: 131K tokens versus 33K.
  • Qwen2.5 32B Instruct has downloadable open weights; the other is API-only.

Side by side

Qwen Max and Qwen2.5 32B Instruct specifications
Qwen MaxQwen2.5 32B Instruct
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.730.1
Released2024-04-032024-09
WeightsProprietaryOpen
Context window33K131K
Max output8K8K
Input $ / M tokens$1.60$0.70
Output $ / M tokens$6.40$2.80
Results tracked237

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 32B Instruct leads

Qwen Max: 30.7 (#292), Qwen2.5 32B Instruct: 38.7 (#169)

Coding benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
Aider Polyglot21.8%—
BigCodeBench Instruct—45%
LMArena Coding1288—
BigCodeBench Complete—52.3%

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Qwen2.5 32B Instruct: 19.2 (#266)

Reasoning benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
Chess Puzzles—0%
LMArena Hard Prompts1269—
Epoch Capabilities Index—128.52

Math Qwen Max leads

Qwen Max: 22.3 (#276), Qwen2.5 32B Instruct: 16.2 (#296)

Math benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
OTIS Mock AIME 2024-202516.1%7.4%
MATH Level 567.2%56.1%
LMArena Math1275—
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen Max leads

Qwen Max: 30.3 (#228), Qwen2.5 32B Instruct: 24.9 (#266)

Knowledge benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
GPQA Diamond56.1%46.1%
LMArena Expert1248—

Multilingual Not comparable

Qwen Max: 41.8 (#202), Qwen2.5 32B Instruct: —

Multilingual benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
LMArena Non-English1263—
LMArena Chinese1254—
LMArena French1330—
LMArena German1254—
LMArena Japanese1205—
LMArena Korean1142—
LMArena Russian1274—
LMArena Spanish1290—

Instruction Following Not comparable

Qwen Max: 66.5 (#208), Qwen2.5 32B Instruct: —

Instruction Following benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
LMArena Instruction Following1262—

Long Context Not comparable

Qwen Max: 39.4 (#180), Qwen2.5 32B Instruct: —

Long Context benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
Fiction.LiveBench66.7%—
LMArena Longer Query1288—

Writing & Preference Not comparable

Qwen Max: 47.8 (#205), Qwen2.5 32B Instruct: —

Writing & Preference benchmarks
BenchmarkQwen MaxQwen2.5 32B Instruct
LMArena Text1282—
LMArena Creative Writing1248—
LMArena Multi-Turn1277—

Frequently asked questions

Is Qwen Max better than Qwen2.5 32B Instruct?

Qwen Max is the stronger model overall, scoring 34.7 to 30.1 on the Noometry Index. Qwen2.5 32B Instruct costs 2.3× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Which is cheaper, Qwen Max or Qwen2.5 32B Instruct?

Qwen2.5 32B Instruct is cheaper. It lists at $0.70 per million input tokens and $2.80 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Qwen Max or Qwen2.5 32B Instruct better for coding?

Qwen2.5 32B Instruct scores higher on coding benchmarks: 38.7 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Qwen2.5 32B Instruct does, with 131K tokens against 33K.

How many benchmarks do Qwen Max and Qwen2.5 32B Instruct share?

3 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen2.5 32B Instruct has 7.

Related comparisons

Go deeper