Model comparison

Qwen3.5 Plus vs Step 3

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 40.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • The widest gap is in math, where Qwen3.5 Plus leads 49.6 to 37.6.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

Qwen3.5 Plus and Step 3 specifications
Qwen3.5 PlusStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index42.940.5
Released2026-02-16—
WeightsProprietaryOpen
Context window1M—
Max output66K—
Input $ / M tokens$0.40—
Output $ / M tokens$2.40—
Results tracked1517

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Qwen3.5 Plus: —, Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.5 PlusStep 3
LMArena Coding—1367
ALE-Bench621.92—

Agentic & Tool Use Not comparable

Qwen3.5 Plus: —, Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5 PlusStep 3
Vending-Bench 20.54—

Reasoning Qwen3.5 Plus leads

Qwen3.5 Plus: 32.8 (#74), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.5 PlusStep 3
Kagi LLM Benchmark—62.3%
Chess Puzzles22%—
LMArena Hard Prompts—1355
Mystery Game Puzzles17%—
DTBench80.5%—
LMCA36.4%—
Epoch Capabilities Index146.78—

Math Qwen3.5 Plus leads

Qwen3.5 Plus: 49.6 (#61), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.5 PlusStep 3
OTIS Mock AIME 2024-202586.7%—
LMArena Math—1366
FrontierMath (Feb 2025 set)21%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Qwen3.5 Plus leads

Qwen3.5 Plus: 46.0 (#83), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.5 PlusStep 3
GPQA Diamond84.8%—
SimpleQA Verified25.4%—
Vectara Hallucination Rate10.7%—
LMArena Expert—1333

Multimodal Not comparable

Qwen3.5 Plus: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.5 PlusStep 3
LMArena Vision—1177

Multilingual Not comparable

Qwen3.5 Plus: —, Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.5 PlusStep 3
LMArena Non-English—1327
LMArena Chinese—1397
LMArena German—1371
LMArena Korean—1269
LMArena Russian—1331
LMArena Spanish—1371

Instruction Following Not comparable

Qwen3.5 Plus: —, Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.5 PlusStep 3
LMArena Instruction Following—1332

Long Context Qwen3.5 Plus leads

Qwen3.5 Plus: 43.0 (#113), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.5 PlusStep 3
CL-bench19.8%—
CL-bench Life12.4%—
LMArena Longer Query—1326

Writing & Preference Not comparable

Qwen3.5 Plus: —, Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.5 PlusStep 3
LMArena Text—1350
LMArena Creative Writing—1321
LMArena Multi-Turn—1341

Frequently asked questions

Is Qwen3.5 Plus better than Step 3?

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 40.5 on the Noometry Index.

How many benchmarks do Qwen3.5 Plus and Step 3 share?

0 benchmarks have published results for both models. Qwen3.5 Plus has 15 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper