Model comparison

Qwen Turbo vs Step 3.5 Flash

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 27.1 on the Noometry Index. Qwen Turbo costs 1.7× less per token, which makes it the better buy when Step 3.5 Flash's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • The widest gap is in math, where Step 3.5 Flash leads 42.6 to 15.3.
  • Qwen Turbo is cheaper at $0.05 / $0.20 per million input/output tokens, against $0.10 / $0.30 for Step 3.5 Flash.
  • Qwen Turbo accepts more context: 1M tokens versus 256K.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

Qwen Turbo and Step 3.5 Flash specifications
Qwen TurboStep 3.5 Flash
ProviderAlibaba (Qwen)StepFun
Noometry Index27.142.3
Released2024-11-012026-01-29
WeightsProprietaryOpen
Context window1M256K
Max output16K256K
Input $ / M tokens$0.05$0.10
Output $ / M tokens$0.20$0.30
Results tracked319

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Qwen Turbo: —, Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkQwen TurboStep 3.5 Flash
LMArena Coding—1436

Reasoning Not comparable

Qwen Turbo: —, Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkQwen TurboStep 3.5 Flash
NYT Connections (extended)—28.4%
LMArena Hard Prompts—1411

Math Step 3.5 Flash leads

Qwen Turbo: 15.3 (#297), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkQwen TurboStep 3.5 Flash
MathArena Final-Answer Competitions—66.8%
OTIS Mock AIME 2024-20256.1%—
LMArena Math—1408
MATH Level 556.2%—

Knowledge Step 3.5 Flash leads

Qwen Turbo: 22.2 (#272), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkQwen TurboStep 3.5 Flash
GPQA Diamond41.8%—
LMArena Expert—1421

Multilingual Not comparable

Qwen Turbo: —, Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkQwen TurboStep 3.5 Flash
LMArena Non-English—1385
LMArena Chinese—1447
LMArena French—1421
LMArena German—1405
LMArena Japanese—1354
LMArena Korean—1352
LMArena Russian—1385
LMArena Spanish—1419

Instruction Following Not comparable

Qwen Turbo: —, Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkQwen TurboStep 3.5 Flash
LMArena Instruction Following—1385

Long Context Not comparable

Qwen Turbo: —, Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkQwen TurboStep 3.5 Flash
LMArena Longer Query—1402

Writing & Preference Not comparable

Qwen Turbo: —, Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkQwen TurboStep 3.5 Flash
LMArena Text—1403
LMArena Creative Writing—1357
LMArena Multi-Turn—1405

Frequently asked questions

Is Qwen Turbo better than Step 3.5 Flash?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 27.1 on the Noometry Index. Qwen Turbo costs 1.7× less per token, which makes it the better buy when Step 3.5 Flash's lead doesn't matter for your workload.

Which is cheaper, Qwen Turbo or Step 3.5 Flash?

Qwen Turbo is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; Step 3.5 Flash lists at $0.10 and $0.30.

Which has the bigger context window?

Qwen Turbo does, with 1M tokens against 256K.

How many benchmarks do Qwen Turbo and Step 3.5 Flash share?

0 benchmarks have published results for both models. Qwen Turbo has 3 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper