Model comparison

Qwen3.6 Max Preview vs Step 3.7 Flash

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 37.3 on the Noometry Index. Step 3.7 Flash costs 7.0× less per token, which makes it the better buy when Qwen3.6 Max Preview's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Step 3.7 Flash StepFun

37.3

Rank #207 Reported

Summary

  • They share 1 benchmark with published results for both. Qwen3.6 Max Preview scores higher in 3 categories and Step 3.7 Flash in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Max Preview leads 41.7 to 21.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 74.1% for Qwen3.6 Max Preview and 39.7% for Step 3.7 Flash.
  • Step 3.7 Flash is cheaper at $0.18 / $1.11 per million input/output tokens, against $1.30 / $7.80 for Qwen3.6 Max Preview.
  • Qwen3.6 Max Preview accepts more context: 262K tokens versus 256K.
  • Step 3.7 Flash has downloadable open weights; the other is API-only.

Side by side

Qwen3.6 Max Preview and Step 3.7 Flash specifications
Qwen3.6 Max PreviewStep 3.7 Flash
ProviderAlibaba (Qwen)StepFun
Noometry Index51.537.3
Released2026-04-202026-05-29
WeightsProprietaryOpen
Context window262K256K
Max output66K256K
Input $ / M tokens$1.30$0.18
Output $ / M tokens$7.80$1.11
Results tracked295

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 48.7 (#54), Step 3.7 Flash: 40.0 (#150)

Coding benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
SWE-bench Verified76.7%—
LMArena WebDev1482—
SciCode—40%
LMArena Coding1471—
ALE-Bench—694.12

Agentic & Tool Use Not comparable

Qwen3.6 Max Preview: —, Step 3.7 Flash: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
Vending-Bench 24,254—

Reasoning Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 41.7 (#53), Step 3.7 Flash: 21.6 (#219)

Reasoning benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
NYT Connections (extended)74.1%39.7%
SimpleBench63%—
CritPt—2.3%
Chess Puzzles20%—
LMArena Hard Prompts1457—
Mystery Game Puzzles19%—
DTBench87.2%—
LMCA42.5%—
Epoch Capabilities Index149.24—

Math Qwen3.6 Max Preview leads

Qwen3.6 Max Preview: 54.1 (#46), Step 3.7 Flash: 42.9 (#82)

Math benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
MathArena Final-Answer Competitions—68.5%
OTIS Mock AIME 2024-202591.1%—
LMArena Math1465—
FrontierMath (Feb 2025 set)23.1%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Not comparable

Qwen3.6 Max Preview: 57.6 (#39), Step 3.7 Flash: —

Knowledge benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
GPQA Diamond87.4%—
SimpleQA Verified52%—
LMArena Expert1478—

Multilingual Not comparable

Qwen3.6 Max Preview: 54.2 (#48), Step 3.7 Flash: —

Multilingual benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
LMArena Non-English1437—
LMArena Chinese1487—
LMArena French1449—
LMArena Russian1445—
LMArena Spanish1454—

Instruction Following Not comparable

Qwen3.6 Max Preview: 75.7 (#55), Step 3.7 Flash: —

Instruction Following benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
LMArena Instruction Following1438—

Long Context Not comparable

Qwen3.6 Max Preview: 44.6 (#61), Step 3.7 Flash: —

Long Context benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
LMArena Longer Query1457—

Writing & Preference Not comparable

Qwen3.6 Max Preview: 63.8 (#60), Step 3.7 Flash: —

Writing & Preference benchmarks
BenchmarkQwen3.6 Max PreviewStep 3.7 Flash
LMArena Text1447—
LMArena Creative Writing1435—
LMArena Multi-Turn1456—

Frequently asked questions

Is Qwen3.6 Max Preview better than Step 3.7 Flash?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 37.3 on the Noometry Index. Step 3.7 Flash costs 7.0× less per token, which makes it the better buy when Qwen3.6 Max Preview's lead doesn't matter for your workload.

Which is cheaper, Qwen3.6 Max Preview or Step 3.7 Flash?

Step 3.7 Flash is cheaper. It lists at $0.18 per million input tokens and $1.11 per million output tokens; Qwen3.6 Max Preview lists at $1.30 and $7.80.

Is Qwen3.6 Max Preview or Step 3.7 Flash better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 40.0 in the Noometry coding category.

Which has the bigger context window?

Qwen3.6 Max Preview does, with 262K tokens against 256K.

How many benchmarks do Qwen3.6 Max Preview and Step 3.7 Flash share?

1 benchmark has published results for both models. Qwen3.6 Max Preview has 29 scored results on Noometry and Step 3.7 Flash has 5.

Related comparisons

Go deeper