Model comparison

Qwen3.5 Max Preview vs Step 3

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 40.5 on the Noometry Index.

Last verified . 15 shared benchmarks.

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Qwen3.5 Max Preview scores higher in 8 categories and Step 3 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 54.3.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

Qwen3.5 Max Preview and Step 3 specifications
Qwen3.5 Max PreviewStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index45.340.5
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1717

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 44.0 (#77), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Coding14871367

Reasoning Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 30.8 (#84), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Hard Prompts14831355
Kagi LLM Benchmark—62.3%

Math Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 40.1 (#94), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Math14741366

Knowledge Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 41.8 (#107), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Expert14891333

Multimodal Not comparable

Qwen3.5 Max Preview: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Vision—1177

Multilingual Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 56.2 (#22), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Non-English14651327
LMArena Chinese15341397
LMArena German14871371
LMArena Korean14381269
LMArena Russian14711331
LMArena Spanish14701371
LMArena French1484—
LMArena Japanese1495—

Instruction Following Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 77.0 (#31), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Instruction Following14671332

Long Context Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 45.2 (#45), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Longer Query14761326

Writing & Preference Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 66.0 (#41), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.5 Max PreviewStep 3
LMArena Text14701350
LMArena Creative Writing14641321
LMArena Multi-Turn14781341

Frequently asked questions

Is Qwen3.5 Max Preview better than Step 3?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 40.5 on the Noometry Index.

Is Qwen3.5 Max Preview or Step 3 better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 40.1 in the Noometry coding category.

How many benchmarks do Qwen3.5 Max Preview and Step 3 share?

15 benchmarks have published results for both models. Qwen3.5 Max Preview has 17 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper