Model comparison

Qwen3.5-Flash vs Step 3

Qwen3.5-Flash is the stronger model overall, scoring 42.5 to 40.5 on the Noometry Index.

Last verified . 15 shared benchmarks.

Qwen3.5-Flash Alibaba (Qwen)

42.5

Rank #112 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Qwen3.5-Flash scores higher in 6 categories and Step 3 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.5-Flash leads 43.2 to 36.8.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

Qwen3.5-Flash and Step 3 specifications
Qwen3.5-FlashStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index42.540.5
Released2026-02-23—
WeightsProprietaryOpen
Context window1M—
Max output66K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.40—
Results tracked3217

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

Qwen3.5-Flash: 34.2 (#242), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Coding14121367
LMArena WebDev1244—
ALE-Bench221.8—

Agentic & Tool Use Not comparable

Qwen3.5-Flash: —, Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5-FlashStep 3
Vending-Bench 2462.69—

Reasoning Qwen3.5-Flash leads

Qwen3.5-Flash: 33.7 (#72), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Hard Prompts14031355
Kagi LLM Benchmark—62.3%
Chess Puzzles21%—
Mystery Game Puzzles20%—
DTBench82.9%—
LMCA29.1%—
Epoch Capabilities Index143.98—

Math Too close to call

Qwen3.5-Flash: 37.4 (#158), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Math14071366
FrontierMath (Tiers 1-3)18.2%—
OTIS Mock AIME 2024-202584.4%—
FrontierMath (Feb 2025 set)6.2%—
FrontierMath Tier 4 (v1)0%—

Knowledge Qwen3.5-Flash leads

Qwen3.5-Flash: 43.2 (#93), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Expert14071333
GPQA Diamond82.3%—
SimpleQA Verified20.3%—
Vectara Hallucination Rate10.5%—

Multimodal Not comparable

Qwen3.5-Flash: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Vision—1177

Multilingual Qwen3.5-Flash leads

Qwen3.5-Flash: 50.5 (#121), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Non-English13851327
LMArena Chinese14461397
LMArena German13901371
LMArena Korean13441269
LMArena Russian13791331
LMArena Spanish14001371
LMArena French1412—
LMArena Japanese1368—

Instruction Following Qwen3.5-Flash leads

Qwen3.5-Flash: 72.6 (#139), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Instruction Following13741332

Long Context Qwen3.5-Flash leads

Qwen3.5-Flash: 42.4 (#124), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Longer Query13921326

Writing & Preference Qwen3.5-Flash leads

Qwen3.5-Flash: 57.9 (#122), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.5-FlashStep 3
LMArena Text13971350
LMArena Creative Writing13431321
LMArena Multi-Turn13931341

Frequently asked questions

Is Qwen3.5-Flash better than Step 3?

Qwen3.5-Flash is the stronger model overall, scoring 42.5 to 40.5 on the Noometry Index.

Is Qwen3.5-Flash or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 34.2 in the Noometry coding category.

How many benchmarks do Qwen3.5-Flash and Step 3 share?

15 benchmarks have published results for both models. Qwen3.5-Flash has 32 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper