Model comparison

Qwen3.5 27B vs Step 3

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 40.5 on the Noometry Index.

Last verified . 16 shared benchmarks.

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Qwen3.5 27B scores higher in 7 categories and Step 3 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 27B leads 59.3 to 54.3.

Side by side

Qwen3.5 27B and Step 3 specifications
Qwen3.5 27BStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index41.940.5
Released2026-02-23—
WeightsOpenOpen
Context window262K—
Max output66K—
Input $ / M tokens$0.30—
Output $ / M tokens$2.40—
Results tracked2817

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

Qwen3.5 27B: 38.9 (#168), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Coding14271367
LMArena WebDev1358—
WeirdML39.5%—
ALE-Bench349.45—

Agentic & Tool Use Not comparable

Qwen3.5 27B: —, Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5 27BStep 3
Vending-Bench 2201.98—

Reasoning Too close to call

Qwen3.5 27B: 27.5 (#117), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Hard Prompts14141355
Kagi LLM Benchmark—62.3%
NYT Connections (extended)47.9%—
Thematic Generalization45.5%—
DTBench82.4%—
LMCA34%—

Math Qwen3.5 27B leads

Qwen3.5 27B: 38.8 (#127), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Math14291366
MathArena Final-Answer Competitions56.7%—

Knowledge Qwen3.5 27B leads

Qwen3.5 27B: 38.0 (#150), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Expert14281333
Vectara Hallucination Rate12.1%—

Multimodal Qwen3.5 27B leads

Qwen3.5 27B: 39.4 (#59), Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Vision12411177

Multilingual Qwen3.5 27B leads

Qwen3.5 27B: 50.8 (#115), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Non-English13901327
LMArena Chinese14781397
LMArena German13931371
LMArena Korean13581269
LMArena Russian13901331
LMArena Spanish14071371
LMArena French1410—
LMArena Japanese1345—

Instruction Following Qwen3.5 27B leads

Qwen3.5 27B: 73.5 (#119), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Instruction Following13931332

Long Context Qwen3.5 27B leads

Qwen3.5 27B: 43.1 (#106), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Longer Query14131326

Writing & Preference Qwen3.5 27B leads

Qwen3.5 27B: 59.3 (#111), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.5 27BStep 3
LMArena Text14091350
LMArena Creative Writing13621321
LMArena Multi-Turn14101341

Frequently asked questions

Is Qwen3.5 27B better than Step 3?

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 40.5 on the Noometry Index.

Is Qwen3.5 27B or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 38.9 in the Noometry coding category.

How many benchmarks do Qwen3.5 27B and Step 3 share?

16 benchmarks have published results for both models. Qwen3.5 27B has 28 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper