Model comparison

Mistral Large 3 vs Step 5 Preview

Step 5 Preview is the stronger model overall, scoring 47.9 to 39.1 on the Noometry Index. Mistral Large 3 costs 3.8× less per token, which makes it the better buy when Step 5 Preview's lead doesn't matter for your workload.

Last verified . 15 shared benchmarks.

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Step 5 Preview StepFun

47.9

Rank #58 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mistral Large 3 scores higher in 0 categories and Step 5 Preview in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Step 5 Preview leads 40.0 to 15.2.
  • Mistral Large 3 is cheaper at $0.25 / $0.75 per million input/output tokens, against $1 / $2.70 for Step 5 Preview.
  • Step 5 Preview accepts more context: 1.02M tokens versus 262K.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Mistral Large 3 and Step 5 Preview specifications
Mistral Large 3Step 5 Preview
ProviderMistral AIStepFun
Noometry Index39.147.9
Released2025-12-022026-09-16
WeightsOpenProprietary
Context window262K1.02M
Max output8K66K
Input $ / M tokens$0.25$1
Output $ / M tokens$0.75$2.70
Results tracked2418

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 5 Preview leads

Mistral Large 3: 34.4 (#237), Step 5 Preview: 51.8 (#36)

Coding benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena WebDev12301564
LMArena Coding14481480
SciCode—58.9%

Reasoning Step 5 Preview leads

Mistral Large 3: 15.2 (#319), Step 5 Preview: 40.0 (#57)

Reasoning benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Hard Prompts14291465
Kagi LLM Benchmark50.9%—
NYT Connections (extended)7.5%—
CritPt—20.9%
Thematic Generalization23%—

Math Step 5 Preview leads

Mistral Large 3: 38.7 (#129), Step 5 Preview: 46.1 (#72)

Math benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Math14141470
ProofBench—42%

Knowledge Step 5 Preview leads

Mistral Large 3: 36.0 (#177), Step 5 Preview: 41.2 (#112)

Knowledge benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Expert14211470
Vectara Hallucination Rate14.5%—

Multimodal Step 5 Preview leads

Mistral Large 3: 38.2 (#66), Step 5 Preview: 41.0 (#41)

Multimodal benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Vision12211267

Multilingual Step 5 Preview leads

Mistral Large 3: 52.5 (#84), Step 5 Preview: 53.8 (#54)

Multilingual benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Non-English14131432
LMArena Chinese14471519
LMArena Russian14111436
LMArena Spanish14401439
LMArena French1455—
LMArena German1437—
LMArena Japanese1394—
LMArena Korean1384—

Instruction Following Step 5 Preview leads

Mistral Large 3: 74.0 (#108), Step 5 Preview: 76.0 (#50)

Instruction Following benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Instruction Following14031444

Long Context Step 5 Preview leads

Mistral Large 3: 43.1 (#105), Step 5 Preview: 44.5 (#64)

Long Context benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Longer Query14131455

Writing & Preference Step 5 Preview leads

Mistral Large 3: 60.0 (#101), Step 5 Preview: 62.8 (#71)

Writing & Preference benchmarks
BenchmarkMistral Large 3Step 5 Preview
LMArena Text14281442
LMArena Creative Writing13861410
LMArena Multi-Turn14291447
EQ-Bench Creative Writing1412—

Frequently asked questions

Is Mistral Large 3 better than Step 5 Preview?

Step 5 Preview is the stronger model overall, scoring 47.9 to 39.1 on the Noometry Index. Mistral Large 3 costs 3.8× less per token, which makes it the better buy when Step 5 Preview's lead doesn't matter for your workload.

Which is cheaper, Mistral Large 3 or Step 5 Preview?

Mistral Large 3 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Step 5 Preview lists at $1 and $2.70.

Is Mistral Large 3 or Step 5 Preview better for coding?

Step 5 Preview scores higher on coding benchmarks: 51.8 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

Step 5 Preview does, with 1.02M tokens against 262K.

How many benchmarks do Mistral Large 3 and Step 5 Preview share?

15 benchmarks have published results for both models. Mistral Large 3 has 24 scored results on Noometry and Step 5 Preview has 18.

Related comparisons

Go deeper