Model comparison

Qwen Max vs Step 5 Preview

Step 5 Preview is the stronger model overall, scoring 47.9 to 34.7 on the Noometry Index.

Last verified . 13 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Step 5 Preview StepFun

47.9

Rank #58 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Qwen Max scores higher in 0 categories and Step 5 Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Step 5 Preview leads 46.1 to 22.3.
  • Step 5 Preview is cheaper at $1 / $2.70 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Step 5 Preview accepts more context: 1.02M tokens versus 33K.

Side by side

Qwen Max and Step 5 Preview specifications
Qwen MaxStep 5 Preview
ProviderAlibaba (Qwen)StepFun
Noometry Index34.747.9
Released2024-04-032026-09-16
WeightsProprietaryProprietary
Context window33K1.02M
Max output8K66K
Input $ / M tokens$1.60$1
Output $ / M tokens$6.40$2.70
Results tracked2318

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 5 Preview leads

Qwen Max: 30.7 (#292), Step 5 Preview: 51.8 (#36)

Coding benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Coding12881480
Aider Polyglot21.8%—
LMArena WebDev—1564
SciCode—58.9%

Reasoning Step 5 Preview leads

Qwen Max: 25.1 (#151), Step 5 Preview: 40.0 (#57)

Reasoning benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Hard Prompts12691465
CritPt—20.9%

Math Step 5 Preview leads

Qwen Max: 22.3 (#276), Step 5 Preview: 46.1 (#72)

Math benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Math12751470
OTIS Mock AIME 2024-202516.1%—
ProofBench—42%
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Step 5 Preview leads

Qwen Max: 30.3 (#228), Step 5 Preview: 41.2 (#112)

Knowledge benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Expert12481470
GPQA Diamond56.1%—

Multimodal Not comparable

Qwen Max: —, Step 5 Preview: 41.0 (#41)

Multimodal benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Vision—1267

Multilingual Step 5 Preview leads

Qwen Max: 41.8 (#202), Step 5 Preview: 53.8 (#54)

Multilingual benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Non-English12631432
LMArena Chinese12541519
LMArena Russian12741436
LMArena Spanish12901439
LMArena French1330—
LMArena German1254—
LMArena Japanese1205—
LMArena Korean1142—

Instruction Following Step 5 Preview leads

Qwen Max: 66.5 (#208), Step 5 Preview: 76.0 (#50)

Instruction Following benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Instruction Following12621444

Long Context Step 5 Preview leads

Qwen Max: 39.4 (#180), Step 5 Preview: 44.5 (#64)

Long Context benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Longer Query12881455
Fiction.LiveBench66.7%—

Writing & Preference Step 5 Preview leads

Qwen Max: 47.8 (#205), Step 5 Preview: 62.8 (#71)

Writing & Preference benchmarks
BenchmarkQwen MaxStep 5 Preview
LMArena Text12821442
LMArena Creative Writing12481410
LMArena Multi-Turn12771447

Frequently asked questions

Is Qwen Max better than Step 5 Preview?

Step 5 Preview is the stronger model overall, scoring 47.9 to 34.7 on the Noometry Index.

Which is cheaper, Qwen Max or Step 5 Preview?

Step 5 Preview is cheaper. It lists at $1 per million input tokens and $2.70 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Qwen Max or Step 5 Preview better for coding?

Step 5 Preview scores higher on coding benchmarks: 51.8 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Step 5 Preview does, with 1.02M tokens against 33K.

How many benchmarks do Qwen Max and Step 5 Preview share?

13 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Step 5 Preview has 18.

Related comparisons

Go deeper