Model comparison

GPT-5.4 mini vs Step 3.5 Flash

GPT-5.4 mini is the stronger model overall, scoring 45.0 to 42.3 on the Noometry Index. Step 3.5 Flash costs 11× less per token, which makes it the better buy when GPT-5.4 mini's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

GPT-5.4 mini OpenAI

45.0

Rank #76 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 18 benchmarks with published results for both. GPT-5.4 mini scores higher in 8 categories and Step 3.5 Flash in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.4 mini leads 51.5 to 39.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 61.8% for GPT-5.4 mini and 28.4% for Step 3.5 Flash.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.75 / $4.50 for GPT-5.4 mini.
  • GPT-5.4 mini accepts more context: 400K tokens versus 256K.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 mini and Step 3.5 Flash specifications
GPT-5.4 miniStep 3.5 Flash
ProviderOpenAIStepFun
Noometry Index45.042.3
Released2026-03-172026-01-29
WeightsProprietaryOpen
Context window400K256K
Max output128K256K
Input $ / M tokens$0.75$0.10
Output $ / M tokens$4.50$0.30
Results tracked4619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 mini leads

GPT-5.4 mini: 45.2 (#72), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Coding14381436
FrontierCode27%—
LMArena WebDev1397—
SciCode49.9%—
WeirdML60.3%—
ALE-Bench1,189—

Agentic & Tool Use Not comparable

GPT-5.4 mini: 29.9 (#81), Step 3.5 Flash: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
DeepResearch Bench36.3%—

Reasoning GPT-5.4 mini leads

GPT-5.4 mini: 30.4 (#85), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
NYT Connections (extended)61.8%28.4%
LMArena Hard Prompts14241411
ARC-AGI-218.9%—
Kagi LLM Benchmark37.9%—
ARC-AGI-163.7%—
CritPt10%—
Chess Puzzles24%—
Thematic Generalization61.7%—
Mystery Game Puzzles11%—
DTBench80%—
LMCA40.8%—
Epoch Capabilities Index148.84—
ForecastBench57—

Math GPT-5.4 mini leads

GPT-5.4 mini: 45.5 (#75), Step 3.5 Flash: 42.6 (#84)

Knowledge GPT-5.4 mini leads

GPT-5.4 mini: 51.5 (#67), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Expert14351421
GPQA Diamond86.9%—
SimpleQA Verified29.4%—
Vectara Hallucination Rate5.5%—

Multimodal Not comparable

GPT-5.4 mini: 39.7 (#56), Step 3.5 Flash: —

Multimodal benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Vision1245—

Multilingual GPT-5.4 mini leads

GPT-5.4 mini: 51.9 (#96), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Non-English14051385
LMArena Chinese14461447
LMArena French14401421
LMArena German14091405
LMArena Japanese13741354
LMArena Korean13681352
LMArena Russian14171385
LMArena Spanish14051419

Instruction Following GPT-5.4 mini leads

GPT-5.4 mini: 74.1 (#102), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Instruction Following14051385

Long Context Too close to call

GPT-5.4 mini: 43.0 (#112), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Longer Query14071402

Writing & Preference GPT-5.4 mini leads

GPT-5.4 mini: 64.0 (#58), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkGPT-5.4 miniStep 3.5 Flash
LMArena Text14121403
LMArena Creative Writing13701357
LMArena Multi-Turn14291405
EQ-Bench Creative Writing1665—

Frequently asked questions

Is GPT-5.4 mini better than Step 3.5 Flash?

GPT-5.4 mini is the stronger model overall, scoring 45.0 to 42.3 on the Noometry Index. Step 3.5 Flash costs 11× less per token, which makes it the better buy when GPT-5.4 mini's lead doesn't matter for your workload.

Which is cheaper, GPT-5.4 mini or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; GPT-5.4 mini lists at $0.75 and $4.50.

Is GPT-5.4 mini or Step 3.5 Flash better for coding?

GPT-5.4 mini scores higher on coding benchmarks: 45.2 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 mini does, with 400K tokens against 256K.

How many benchmarks do GPT-5.4 mini and Step 3.5 Flash share?

18 benchmarks have published results for both models. GPT-5.4 mini has 46 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper