Model comparison

DeepSeek-V3.1 vs Step 3.5 Flash

DeepSeek-V3.1 and Step 3.5 Flash score almost the same on the Noometry Index (42.8 vs 42.3), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 5 categories and Step 3.5 Flash in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Step 3.5 Flash leads 42.8 to 36.3.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.25 / $0.95 for DeepSeek-V3.1.
  • Step 3.5 Flash accepts more context: 256K tokens versus 164K.

Side by side

DeepSeek-V3.1 and Step 3.5 Flash specifications
DeepSeek-V3.1Step 3.5 Flash
ProviderDeepSeekStepFun
Noometry Index42.842.3
Released2025-08-212026-01-29
WeightsOpenOpen
Context window164K256K
Max output8K256K
Input $ / M tokens$0.25$0.10
Output $ / M tokens$0.95$0.30
Results tracked2719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.5 Flash leads

DeepSeek-V3.1: 40.3 (#144), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Coding14171436
WeirdML38.4%—

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Hard Prompts14171411
SimpleBench40%—
Kagi LLM Benchmark53.2%—
NYT Connections (extended)—28.4%
DTBench82.7%—
LMCA24.3%—
Epoch Capabilities Index139.92—
ForecastBench58—

Math Step 3.5 Flash leads

DeepSeek-V3.1: 38.9 (#122), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Math14201408
MathArena Final-Answer Competitions—66.8%

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Expert14051421
Vectara Hallucination Rate5.5%—

Multilingual DeepSeek-V3.1 leads

DeepSeek-V3.1: 51.6 (#106), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Non-English14001385
LMArena Chinese14691447
LMArena French14471421
LMArena German14111405
LMArena Japanese13781354
LMArena Korean13371352
LMArena Russian14051385
LMArena Spanish14311419

Instruction Following Too close to call

DeepSeek-V3.1: 73.9 (#110), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Instruction Following14001385

Long Context Step 3.5 Flash leads

DeepSeek-V3.1: 36.3 (#232), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Longer Query14221402
Fiction.LiveBench52.8%—

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Step 3.5 Flash
LMArena Text14201403
LMArena Creative Writing14011357
LMArena Multi-Turn14081405
EQ-Bench Creative Writing1436—

Frequently asked questions

Is DeepSeek-V3.1 better than Step 3.5 Flash?

DeepSeek-V3.1 and Step 3.5 Flash score almost the same on the Noometry Index (42.8 vs 42.3), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.1 or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; DeepSeek-V3.1 lists at $0.25 and $0.95.

Is DeepSeek-V3.1 or Step 3.5 Flash better for coding?

Step 3.5 Flash scores higher on coding benchmarks: 42.4 versus 40.3 in the Noometry coding category.

Which has the bigger context window?

Step 3.5 Flash does, with 256K tokens against 164K.

How many benchmarks do DeepSeek-V3.1 and Step 3.5 Flash share?

17 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper