Model comparison

Mistral Small vs Step 3.5 Flash

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 33.4 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral Small scores higher in 0 categories and Step 3.5 Flash in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Step 3.5 Flash leads 42.6 to 16.4.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.15 / $0.60 for Mistral Small.
  • Mistral Small accepts more context: 262K tokens versus 256K.

Side by side

Mistral Small and Step 3.5 Flash specifications
Mistral SmallStep 3.5 Flash
ProviderMistral AIStepFun
Noometry Index33.442.3
Released2024-02-262026-01-29
WeightsOpenOpen
Context window262K256K
Max output256K256K
Input $ / M tokens$0.15$0.10
Output $ / M tokens$0.60$0.30
Results tracked3919

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.5 Flash leads

Mistral Small: 34.0 (#247), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Coding13621436
SciCode26.5%—
BigCodeBench Instruct36.1%—
LiveBench Coding36.2%—
BigCodeBench Complete46.6%—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Step 3.5 Flash: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallStep 3.5 Flash
Berkeley Function Calling Leaderboard37.1%—

Reasoning Step 3.5 Flash leads

Mistral Small: 19.8 (#250), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Hard Prompts13351411
Kagi LLM Benchmark37.8%—
NYT Connections (extended)—28.4%
CritPt0%—
LiveBench Reasoning44.8%—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
LiveBench44%—

Math Step 3.5 Flash leads

Mistral Small: 16.4 (#293), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Math13411408
MathArena Final-Answer Competitions—66.8%
OTIS Mock AIME 2024-20255.8%—
LiveBench Math39.9%—
MATH Level 546.8%—

Knowledge Step 3.5 Flash leads

Mistral Small: 31.0 (#222), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Expert12911421
GPQA Diamond47.5%—
Vectara Hallucination Rate5.1%—
MMLU68.7%—

Multimodal Not comparable

Mistral Small: 33.5 (#96), Step 3.5 Flash: —

Multimodal benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Vision1142—

Multilingual Step 3.5 Flash leads

Mistral Small: 45.5 (#169), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Non-English13151385
LMArena Chinese13401447
LMArena French13371421
LMArena German13401405
LMArena Japanese12751354
LMArena Korean12591352
LMArena Russian13241385
LMArena Spanish13461419

Instruction Following Step 3.5 Flash leads

Mistral Small: 66.4 (#209), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Instruction Following13101385
LiveBench Instruction Following63.7%—

Long Context Step 3.5 Flash leads

Mistral Small: 40.4 (#156), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Longer Query13271402

Writing & Preference Step 3.5 Flash leads

Mistral Small: 52.5 (#171), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkMistral SmallStep 3.5 Flash
LMArena Text13381403
LMArena Creative Writing13051357
LMArena Multi-Turn13441405
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Step 3.5 Flash?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 33.4 on the Noometry Index.

Which is cheaper, Mistral Small or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Mistral Small lists at $0.15 and $0.60.

Is Mistral Small or Step 3.5 Flash better for coding?

Step 3.5 Flash scores higher on coding benchmarks: 42.4 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 256K.

How many benchmarks do Mistral Small and Step 3.5 Flash share?

17 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper