Model comparison

Qwen3.6 35B-A3B vs Step 2 16k Exp 202412

Step 2 16k Exp 202412 is the stronger model overall, scoring 39.2 to 37.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3.6 35B-A3B Alibaba (Qwen)

37.6

Rank #201 Confirmed

Step 2 16k Exp 202412 StepFun

39.2

Rank #171 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen3.6 35B-A3B leads 51.3 to 35.2.
  • Qwen3.6 35B-A3B has downloadable open weights; the other is API-only.

Side by side

Qwen3.6 35B-A3B and Step 2 16k Exp 202412 specifications
Qwen3.6 35B-A3BStep 2 16k Exp 202412
ProviderAlibaba (Qwen)StepFun
Noometry Index37.639.2
Released2026-04-01—
WeightsOpenProprietary
Context window262K—
Max output66K—
Input $ / M tokens$0.25—
Output $ / M tokens$1.49—
Results tracked1412

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 2 16k Exp 202412 leads

Qwen3.6 35B-A3B: 37.2 (#196), Step 2 16k Exp 202412: 38.5 (#174)

Coding benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
SciCode35.8%—
WeirdML34.5%—
LMArena Coding—1316

Agentic & Tool Use Not comparable

Qwen3.6 35B-A3B: 22.1 (#134), Step 2 16k Exp 202412: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
Terminal-Bench23%—

Reasoning Qwen3.6 35B-A3B leads

Qwen3.6 35B-A3B: 28.0 (#109), Step 2 16k Exp 202412: 25.9 (#142)

Reasoning benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
NYT Connections (extended)41.6%—
CritPt0.3%—
Chess Puzzles26%—
LMArena Hard Prompts—1299
Mystery Game Puzzles22%—
DTBench73.9%—
LMCA29.7%—
Surface Evolver Bench44.4%—
Epoch Capabilities Index143.93—

Math Qwen3.6 35B-A3B leads

Qwen3.6 35B-A3B: 38.9 (#121), Step 2 16k Exp 202412: 36.3 (#169)

Math benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
FrontierMath (Tiers 1-3)20.4%—
OTIS Mock AIME 2024-202586.7%—
LMArena Math—1304

Knowledge Qwen3.6 35B-A3B leads

Qwen3.6 35B-A3B: 51.3 (#68), Step 2 16k Exp 202412: 35.2 (#187)

Knowledge benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
GPQA Diamond84.8%—
LMArena Expert—1279

Multilingual Not comparable

Qwen3.6 35B-A3B: —, Step 2 16k Exp 202412: 43.7 (#181)

Multilingual benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
LMArena Non-English—1290
LMArena Chinese—1331
LMArena Russian—1324

Instruction Following Not comparable

Qwen3.6 35B-A3B: —, Step 2 16k Exp 202412: 67.9 (#192)

Instruction Following benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
LMArena Instruction Following—1287

Long Context Not comparable

Qwen3.6 35B-A3B: —, Step 2 16k Exp 202412: 39.7 (#170)

Long Context benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
LMArena Longer Query—1307

Writing & Preference Not comparable

Qwen3.6 35B-A3B: —, Step 2 16k Exp 202412: 52.3 (#173)

Writing & Preference benchmarks
BenchmarkQwen3.6 35B-A3BStep 2 16k Exp 202412
LMArena Text—1321
LMArena Creative Writing—1328
LMArena Multi-Turn—1292

Frequently asked questions

Is Qwen3.6 35B-A3B better than Step 2 16k Exp 202412?

Step 2 16k Exp 202412 is the stronger model overall, scoring 39.2 to 37.6 on the Noometry Index.

Is Qwen3.6 35B-A3B or Step 2 16k Exp 202412 better for coding?

Step 2 16k Exp 202412 scores higher on coding benchmarks: 38.5 versus 37.2 in the Noometry coding category.

How many benchmarks do Qwen3.6 35B-A3B and Step 2 16k Exp 202412 share?

0 benchmarks have published results for both models. Qwen3.6 35B-A3B has 14 scored results on Noometry and Step 2 16k Exp 202412 has 12.

Related comparisons

Go deeper