Model comparison

Qwen3.6 27B vs Step 3

Qwen3.6 27B is the stronger model overall, scoring 42.2 to 40.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3.6 27B Alibaba (Qwen)

42.2

Rank #117 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen3.6 27B leads 52.4 to 36.8.

Side by side

Qwen3.6 27B and Step 3 specifications
Qwen3.6 27BStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index42.240.5
Released2026-04-22—
WeightsOpenOpen
Context window262K—
Max output66K—
Input $ / M tokens$0.60—
Output $ / M tokens$3.60—
Results tracked1117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

Qwen3.6 27B: 39.1 (#163), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.6 27BStep 3
SciCode37.3%—
LMArena Coding—1367

Reasoning Step 3 leads

Qwen3.6 27B: 25.0 (#153), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.6 27BStep 3
Kagi LLM Benchmark—62.3%
CritPt0.9%—
Chess Puzzles22%—
LMArena Hard Prompts—1355
Mystery Game Puzzles7%—
DTBench78.1%—
LMCA34.5%—
Epoch Capabilities Index146.5—

Math Qwen3.6 27B leads

Qwen3.6 27B: 48.5 (#62), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.6 27BStep 3
FrontierMath (Tiers 1-3)35.1%—
OTIS Mock AIME 2024-202591.1%—
LMArena Math—1366

Knowledge Qwen3.6 27B leads

Qwen3.6 27B: 52.4 (#63), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.6 27BStep 3
GPQA Diamond85.9%—
LMArena Expert—1333

Multimodal Not comparable

Qwen3.6 27B: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.6 27BStep 3
LMArena Vision—1177

Multilingual Not comparable

Qwen3.6 27B: —, Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.6 27BStep 3
LMArena Non-English—1327
LMArena Chinese—1397
LMArena German—1371
LMArena Korean—1269
LMArena Russian—1331
LMArena Spanish—1371

Instruction Following Not comparable

Qwen3.6 27B: —, Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.6 27BStep 3
LMArena Instruction Following—1332

Long Context Not comparable

Qwen3.6 27B: —, Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.6 27BStep 3
LMArena Longer Query—1326

Writing & Preference Step 3 leads

Qwen3.6 27B: 50.3 (#181), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.6 27BStep 3
LMArena Text—1350
LMArena Creative Writing—1321
EQ-Bench 41026—
LMArena Multi-Turn—1341

Frequently asked questions

Is Qwen3.6 27B better than Step 3?

Qwen3.6 27B is the stronger model overall, scoring 42.2 to 40.5 on the Noometry Index.

Is Qwen3.6 27B or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 39.1 in the Noometry coding category.

How many benchmarks do Qwen3.6 27B and Step 3 share?

0 benchmarks have published results for both models. Qwen3.6 27B has 11 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper