Model comparison

Phi-4 Mini vs Step 3

Step 3 is the stronger model overall, scoring 40.5 to 30.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • The widest gap is in coding, where Step 3 leads 40.1 to 28.1.

Side by side

Phi-4 Mini and Step 3 specifications
Phi-4 MiniStep 3
ProviderMicrosoftStepFun
Noometry Index30.940.5
Released2024-12-11—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.075—
Output $ / M tokens$0.30—
Results tracked317

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

Phi-4 Mini: 28.1 (#317), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkPhi-4 MiniStep 3
SciCode10.8%—
LMArena Coding—1367

Reasoning Step 3 leads

Phi-4 Mini: 22.4 (#195), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkPhi-4 MiniStep 3
Kagi LLM Benchmark—62.3%
CritPt0%—
LMArena Hard Prompts—1355

Math Not comparable

Phi-4 Mini: —, Step 3: 37.6 (#148)

Math benchmarks
BenchmarkPhi-4 MiniStep 3
LMArena Math—1366

Knowledge Step 3 leads

Phi-4 Mini: 25.3 (#262), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkPhi-4 MiniStep 3
Vectara Hallucination Rate23.5%—
LMArena Expert—1333

Multimodal Not comparable

Phi-4 Mini: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkPhi-4 MiniStep 3
LMArena Vision—1177

Multilingual Not comparable

Phi-4 Mini: —, Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkPhi-4 MiniStep 3
LMArena Non-English—1327
LMArena Chinese—1397
LMArena German—1371
LMArena Korean—1269
LMArena Russian—1331
LMArena Spanish—1371

Instruction Following Not comparable

Phi-4 Mini: —, Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkPhi-4 MiniStep 3
LMArena Instruction Following—1332

Long Context Not comparable

Phi-4 Mini: —, Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkPhi-4 MiniStep 3
LMArena Longer Query—1326

Writing & Preference Not comparable

Phi-4 Mini: —, Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkPhi-4 MiniStep 3
LMArena Text—1350
LMArena Creative Writing—1321
LMArena Multi-Turn—1341

Frequently asked questions

Is Phi-4 Mini better than Step 3?

Step 3 is the stronger model overall, scoring 40.5 to 30.9 on the Noometry Index.

Is Phi-4 Mini or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 28.1 in the Noometry coding category.

How many benchmarks do Phi-4 Mini and Step 3 share?

0 benchmarks have published results for both models. Phi-4 Mini has 3 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper