Model comparison

Mercury 2.5 vs Step 3.5 Flash

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 33.5 on the Noometry Index. Mercury 2.5 costs 2.2× less per token, which makes it the better buy when Step 3.5 Flash's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • The widest gap is in math, where Step 3.5 Flash leads 42.6 to 23.3.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.10 / $0.30 for Step 3.5 Flash.
  • Mercury 2.5 accepts more context: 260K tokens versus 256K.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Step 3.5 Flash specifications
Mercury 2.5Step 3.5 Flash
ProviderInceptionStepFun
Noometry Index33.542.3
Released2026-09-082026-01-29
WeightsProprietaryOpen
Context window260K256K
Max output66K256K
Input $ / M tokens$0.04$0.10
Output $ / M tokens$0.15$0.30
Results tracked419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.5 Flash leads

Mercury 2.5: 39.5 (#156), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
SciCode38.5%—
LMArena Coding—1436
ALE-Bench301.65—

Reasoning Too close to call

Mercury 2.5: 22.4 (#193), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
NYT Connections (extended)—28.4%
CritPt0%—
LMArena Hard Prompts—1411

Math Step 3.5 Flash leads

Mercury 2.5: 23.3 (#272), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
MathArena Final-Answer Competitions—66.8%
ProofBench3%—
LMArena Math—1408

Knowledge Not comparable

Mercury 2.5: —, Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
LMArena Expert—1421

Multilingual Not comparable

Mercury 2.5: —, Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
LMArena Non-English—1385
LMArena Chinese—1447
LMArena French—1421
LMArena German—1405
LMArena Japanese—1354
LMArena Korean—1352
LMArena Russian—1385
LMArena Spanish—1419

Instruction Following Not comparable

Mercury 2.5: —, Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
LMArena Instruction Following—1385

Long Context Not comparable

Mercury 2.5: —, Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
LMArena Longer Query—1402

Writing & Preference Not comparable

Mercury 2.5: —, Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkMercury 2.5Step 3.5 Flash
LMArena Text—1403
LMArena Creative Writing—1357
LMArena Multi-Turn—1405

Frequently asked questions

Is Mercury 2.5 better than Step 3.5 Flash?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 33.5 on the Noometry Index. Mercury 2.5 costs 2.2× less per token, which makes it the better buy when Step 3.5 Flash's lead doesn't matter for your workload.

Which is cheaper, Mercury 2.5 or Step 3.5 Flash?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Step 3.5 Flash lists at $0.10 and $0.30.

Is Mercury 2.5 or Step 3.5 Flash better for coding?

Step 3.5 Flash scores higher on coding benchmarks: 42.4 versus 39.5 in the Noometry coding category.

Which has the bigger context window?

Mercury 2.5 does, with 260K tokens against 256K.

How many benchmarks do Mercury 2.5 and Step 3.5 Flash share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper