Model comparison

Qwen3.8 27B vs Step 3

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 40.5 on the Noometry Index.

Last verified . 16 shared benchmarks.

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Qwen3.8 27B scores higher in 8 categories and Step 3 in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.8 27B leads 41.0 to 28.4.

Side by side

Qwen3.8 27B and Step 3 specifications
Qwen3.8 27BStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index46.040.5
Released2026-08-14—
WeightsOpenOpen
Context window262K—
Max output33K—
Input $ / M tokens$0.99—
Output $ / M tokens$1.49—
Results tracked3117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

Qwen3.8 27B: 50.5 (#44), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Coding14821367
LMArena WebDev1593—
SciCode46.6%—

Agentic & Tool Use Not comparable

Qwen3.8 27B: 32.9 (#57), Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.8 27BStep 3
APEX-Agents47.5%—

Reasoning Qwen3.8 27B leads

Qwen3.8 27B: 41.0 (#54), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Hard Prompts14601355
ARC-AGI-242.4%—
Kagi LLM Benchmark—62.3%
NYT Connections (extended)54.5%—
ARC-AGI-187.5%—
CritPt5.4%—
DTBench88%—
LMCA41.4%—
Surface Evolver Bench45%—
Epoch Capabilities Index149.38—

Math Too close to call

Qwen3.8 27B: 37.1 (#161), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Math14561366
ProofBench16%—

Knowledge Qwen3.8 27B leads

Qwen3.8 27B: 41.6 (#109), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Expert14821333

Multimodal Qwen3.8 27B leads

Qwen3.8 27B: 41.3 (#37), Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Vision12711177

Multilingual Qwen3.8 27B leads

Qwen3.8 27B: 53.7 (#60), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Non-English14301327
LMArena Chinese15041397
LMArena German14381371
LMArena Korean13931269
LMArena Russian14151331
LMArena Spanish14481371
LMArena French1465—
LMArena Japanese1384—

Instruction Following Qwen3.8 27B leads

Qwen3.8 27B: 75.8 (#53), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Instruction Following14391332

Long Context Qwen3.8 27B leads

Qwen3.8 27B: 44.3 (#70), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Longer Query14501326

Writing & Preference Qwen3.8 27B leads

Qwen3.8 27B: 65.8 (#43), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.8 27BStep 3
LMArena Text14411350
LMArena Creative Writing13841321
LMArena Multi-Turn14411341
EQ-Bench Creative Writing1671—

Frequently asked questions

Is Qwen3.8 27B better than Step 3?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 40.5 on the Noometry Index.

Is Qwen3.8 27B or Step 3 better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 40.1 in the Noometry coding category.

How many benchmarks do Qwen3.8 27B and Step 3 share?

16 benchmarks have published results for both models. Qwen3.8 27B has 31 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper