Model comparison

Qwen3.6 Max Preview vs Step 3

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 40.5 on the Noometry Index.

Last verified . 13 shared benchmarks.

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Qwen3.6 Max Preview scores higher in 8 categories and Step 3 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Max Preview leads 57.6 to 36.8.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

Qwen3.6 Max Preview and Step 3 specifications
Qwen3.6 Max PreviewStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index51.540.5
Released2026-04-20—
WeightsProprietaryOpen
Context window262K—
Max output66K—
Input $ / M tokens$1.30—
Output $ / M tokens$7.80—
Results tracked2917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 48.7 (#54), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Coding14711367
SWE-bench Verified76.7%—
LMArena WebDev1482—

Agentic & Tool Use Not comparable

Qwen3.6 Max Preview: —, Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
Vending-Bench 24,254—

Reasoning Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 41.7 (#53), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Hard Prompts14571355
SimpleBench63%—
Kagi LLM Benchmark—62.3%
NYT Connections (extended)74.1%—
Chess Puzzles20%—
Mystery Game Puzzles19%—
DTBench87.2%—
LMCA42.5%—
Epoch Capabilities Index149.24—

Math Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 54.1 (#46), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Math14651366
OTIS Mock AIME 2024-202591.1%—
FrontierMath (Feb 2025 set)23.1%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 57.6 (#39), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Expert14781333
GPQA Diamond87.4%—
SimpleQA Verified52%—

Multimodal Not comparable

Qwen3.6 Max Preview: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Vision—1177

Multilingual Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 54.2 (#48), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Non-English14371327
LMArena Chinese14871397
LMArena Russian14451331
LMArena Spanish14541371
LMArena French1449—
LMArena German—1371
LMArena Korean—1269

Instruction Following Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 75.7 (#55), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Instruction Following14381332

Long Context Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 44.6 (#61), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Longer Query14571326

Writing & Preference Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 63.8 (#60), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.6 Max PreviewStep 3
LMArena Text14471350
LMArena Creative Writing14351321
LMArena Multi-Turn14561341

Frequently asked questions

Is Qwen3.6 Max Preview better than Step 3?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 40.5 on the Noometry Index.

Is Qwen3.6 Max Preview or Step 3 better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 40.1 in the Noometry coding category.

How many benchmarks do Qwen3.6 Max Preview and Step 3 share?

13 benchmarks have published results for both models. Qwen3.6 Max Preview has 29 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper