Model comparison

Magistral Medium vs Step 3

Step 3 is the stronger model overall, scoring 40.5 to 35.2 on the Noometry Index.

Last verified . 16 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Magistral Medium scores higher in 0 categories and Step 3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Step 3 leads 28.4 to 8.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 16.2% for Magistral Medium and 62.3% for Step 3.

Side by side

Magistral Medium and Step 3 specifications
Magistral MediumStep 3
ProviderMistral AIStepFun
Noometry Index35.240.5
Released2025-03-17—
WeightsOpenOpen
Context window262K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$5—
Results tracked2217

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

Magistral Medium: 39.1 (#161), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkMagistral MediumStep 3
LMArena Coding13191367
SciCode39.2%—

Reasoning Step 3 leads

Magistral Medium: 8.6 (#348), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkMagistral MediumStep 3
Kagi LLM Benchmark16.2%62.3%
LMArena Hard Prompts12671355
ARC-AGI-20%—
ARC-AGI-16.1%—
CritPt0.3%—

Math Step 3 leads

Magistral Medium: 35.1 (#189), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkMagistral MediumStep 3
LMArena Math12501366

Knowledge Step 3 leads

Magistral Medium: 33.5 (#202), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkMagistral MediumStep 3
LMArena Expert12231333

Multimodal Not comparable

Magistral Medium: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkMagistral MediumStep 3
LMArena Vision—1177

Multilingual Step 3 leads

Magistral Medium: 39.6 (#224), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkMagistral MediumStep 3
LMArena Non-English12321327
LMArena Chinese12271397
LMArena German12481371
LMArena Korean11251269
LMArena Russian12241331
LMArena Spanish12711371
LMArena French1267—
LMArena Japanese1175—

Instruction Following Step 3 leads

Magistral Medium: 66.0 (#211), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkMagistral MediumStep 3
LMArena Instruction Following12541332

Long Context Step 3 leads

Magistral Medium: 39.3 (#183), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkMagistral MediumStep 3
LMArena Longer Query12951326

Writing & Preference Step 3 leads

Magistral Medium: 46.3 (#219), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkMagistral MediumStep 3
LMArena Text12551350
LMArena Creative Writing12451321
LMArena Multi-Turn12751341

Frequently asked questions

Is Magistral Medium better than Step 3?

Step 3 is the stronger model overall, scoring 40.5 to 35.2 on the Noometry Index.

Is Magistral Medium or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 39.1 in the Noometry coding category.

How many benchmarks do Magistral Medium and Step 3 share?

16 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper