Model comparison

Qwen3.5 397B-A17B vs Step 3.5 Flash

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 42.3 on the Noometry Index. Step 3.5 Flash costs 9.0× less per token, which makes it the better buy when Qwen3.5 397B-A17B's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

Qwen3.5 397B-A17B Alibaba (Qwen)

46.0

Rank #67 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Qwen3.5 397B-A17B scores higher in 7 categories and Step 3.5 Flash in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.5 397B-A17B leads 53.3 to 39.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 58.9% for Qwen3.5 397B-A17B and 28.4% for Step 3.5 Flash.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.60 / $3.60 for Qwen3.5 397B-A17B.
  • Qwen3.5 397B-A17B accepts more context: 262K tokens versus 256K.

Side by side

Qwen3.5 397B-A17B and Step 3.5 Flash specifications
Qwen3.5 397B-A17BStep 3.5 Flash
ProviderAlibaba (Qwen)StepFun
Noometry Index46.042.3
Released2026-02-012026-01-29
WeightsOpenOpen
Context window262K256K
Max output66K256K
Input $ / M tokens$0.60$0.10
Output $ / M tokens$3.60$0.30
Results tracked3619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Qwen3.5 397B-A17B: 42.0 (#114), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Coding14651436
LMArena WebDev1400—

Agentic & Tool Use Not comparable

Qwen3.5 397B-A17B: 33.3 (#53), Step 3.5 Flash: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
APEX-Agents24.9%—
τ²-bench Airline81.5%—
τ²-bench Banking9.8%—
τ²-bench Retail84.4%—
τ²-bench Telecom97.8%—

Reasoning Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 34.5 (#70), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
NYT Connections (extended)58.9%28.4%
LMArena Hard Prompts14481411
Kagi LLM Benchmark73.7%—
Chess Puzzles13%—
Thematic Generalization65.1%—
Mystery Game Puzzles18%—
DTBench87.5%—
LMCA37.9%—
Epoch Capabilities Index146.65—

Math Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 46.1 (#73), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Math14541408
FrontierMath (Tiers 1-3)31.2%—
MathArena Final-Answer Competitions—66.8%
OTIS Mock AIME 2024-202588.9%—

Knowledge Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 53.3 (#58), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Expert14621421
GPQA Diamond86.4%—

Multimodal Not comparable

Qwen3.5 397B-A17B: 40.7 (#44), Step 3.5 Flash: —

Multimodal benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Vision1263—

Multilingual Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 53.7 (#59), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Non-English14301385
LMArena Chinese15001447
LMArena French14611421
LMArena German14471405
LMArena Japanese14261354
LMArena Korean13841352
LMArena Russian14291385
LMArena Spanish14411419

Instruction Following Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 75.0 (#77), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Instruction Following14241385

Long Context Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 44.1 (#74), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Longer Query14421402

Writing & Preference Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 62.3 (#79), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkQwen3.5 397B-A17BStep 3.5 Flash
LMArena Text14381403
LMArena Creative Writing14011357
LMArena Multi-Turn14461405
EQ-Bench Creative Writing1478—

Frequently asked questions

Is Qwen3.5 397B-A17B better than Step 3.5 Flash?

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 42.3 on the Noometry Index. Step 3.5 Flash costs 9.0× less per token, which makes it the better buy when Qwen3.5 397B-A17B's lead doesn't matter for your workload.

Which is cheaper, Qwen3.5 397B-A17B or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Qwen3.5 397B-A17B lists at $0.60 and $3.60.

Is Qwen3.5 397B-A17B or Step 3.5 Flash better for coding?

They score almost the same on coding (42.0 vs 42.4); test both on your own repository before choosing.

Which has the bigger context window?

Qwen3.5 397B-A17B does, with 262K tokens against 256K.

How many benchmarks do Qwen3.5 397B-A17B and Step 3.5 Flash share?

18 benchmarks have published results for both models. Qwen3.5 397B-A17B has 36 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper