Model comparison

Qwen3.5 397B-A17B vs Step 3

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 40.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

Qwen3.5 397B-A17B Alibaba (Qwen)

46.0

Rank #67 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Qwen3.5 397B-A17B scores higher in 9 categories and Step 3 in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.5 397B-A17B leads 53.3 to 36.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 73.7% for Qwen3.5 397B-A17B and 62.3% for Step 3.

Side by side

Qwen3.5 397B-A17B and Step 3 specifications
Qwen3.5 397B-A17BStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index46.040.5
Released2026-02-01—
WeightsOpenOpen
Context window262K—
Max output66K—
Input $ / M tokens$0.60—
Output $ / M tokens$3.60—
Results tracked3617

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 42.0 (#114), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Coding14651367
LMArena WebDev1400—

Agentic & Tool Use Not comparable

Qwen3.5 397B-A17B: 33.3 (#53), Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
APEX-Agents24.9%—
τ²-bench Airline81.5%—
τ²-bench Banking9.8%—
τ²-bench Retail84.4%—
τ²-bench Telecom97.8%—

Reasoning Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 34.5 (#70), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
Kagi LLM Benchmark73.7%62.3%
LMArena Hard Prompts14481355
NYT Connections (extended)58.9%—
Chess Puzzles13%—
Thematic Generalization65.1%—
Mystery Game Puzzles18%—
DTBench87.5%—
LMCA37.9%—
Epoch Capabilities Index146.65—

Math Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 46.1 (#73), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Math14541366
FrontierMath (Tiers 1-3)31.2%—
OTIS Mock AIME 2024-202588.9%—

Knowledge Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 53.3 (#58), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Expert14621333
GPQA Diamond86.4%—

Multimodal Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 40.7 (#44), Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Vision12631177

Multilingual Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 53.7 (#59), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Non-English14301327
LMArena Chinese15001397
LMArena German14471371
LMArena Korean13841269
LMArena Russian14291331
LMArena Spanish14411371
LMArena French1461—
LMArena Japanese1426—

Instruction Following Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 75.0 (#77), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Instruction Following14241332

Long Context Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 44.1 (#74), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Longer Query14421326

Writing & Preference Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 62.3 (#79), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.5 397B-A17BStep 3
LMArena Text14381350
LMArena Creative Writing14011321
LMArena Multi-Turn14461341
EQ-Bench Creative Writing1478—

Frequently asked questions

Is Qwen3.5 397B-A17B better than Step 3?

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 40.5 on the Noometry Index.

Is Qwen3.5 397B-A17B or Step 3 better for coding?

Qwen3.5 397B-A17B scores higher on coding benchmarks: 42.0 versus 40.1 in the Noometry coding category.

How many benchmarks do Qwen3.5 397B-A17B and Step 3 share?

17 benchmarks have published results for both models. Qwen3.5 397B-A17B has 36 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper