Model comparison

Magistral Small vs Phi-4

Magistral Small and Phi-4 score almost the same on the Noometry Index (30.2 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Magistral Small scores higher in 2 categories and Phi-4 in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Phi-4 leads 17.7 to 6.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 30% for Magistral Small and 13.8% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $0.50 / $1.50 for Magistral Small.

Side by side

Magistral Small and Phi-4 specifications
Magistral SmallPhi-4
ProviderMistral AIMicrosoft
Noometry Index30.231.2
Released2025-06-102024-12-11
WeightsOpenOpen
Context window128K128K
Max output40K4K
Input $ / M tokens$0.50$0.07
Output $ / M tokens$1.50$0.14
Results tracked1037

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Magistral Small: 38.4 (#176), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkMagistral SmallPhi-4
SciCode35.2%—
BigCodeBench Instruct—45.5%
LiveBench Coding—30.7%
LMArena Coding—1231
BigCodeBench Complete—55.4%

Agentic & Tool Use Not comparable

Magistral Small: —, Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkMagistral SmallPhi-4
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%

Reasoning Phi-4 leads

Magistral Small: 6.8 (#350), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkMagistral SmallPhi-4
Chess Puzzles3%1%
Epoch Capabilities Index133.19130.42
ARC-AGI-20%—
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
CritPt0.3%—
LiveBench Reasoning—47.8%
LMArena Hard Prompts—1220
DTBench61.3%—
LiveBench Data Analysis—45.2%
LiveBench—41.6%

Math Magistral Small leads

Magistral Small: 26.2 (#261), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkMagistral SmallPhi-4
OTIS Mock AIME 2024-202530%13.8%
LiveBench Math—42%
LMArena Math—1246
MATH Level 5—64.9%

Knowledge Phi-4 leads

Magistral Small: 30.9 (#223), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkMagistral SmallPhi-4
GPQA Diamond56.1%56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%
LMArena Expert—1203
MMLU—84.8%

Multilingual Not comparable

Magistral Small: —, Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkMagistral SmallPhi-4
LMArena Non-English—1197
LMArena Chinese—1212
LMArena French—1224
LMArena German—1222
LMArena Japanese—1158
LMArena Korean—1151
LMArena Russian—1209
LMArena Spanish—1234

Instruction Following Not comparable

Magistral Small: —, Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkMagistral SmallPhi-4
LiveBench Instruction Following—58.4%
LMArena Instruction Following—1201

Long Context Not comparable

Magistral Small: —, Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkMagistral SmallPhi-4
LMArena Longer Query—1217

Writing & Preference Not comparable

Magistral Small: —, Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkMagistral SmallPhi-4
LMArena Text—1217
LMArena Creative Writing—1182
Short-Story Creative Writing—62.6%
LMArena Multi-Turn—1206
LiveBench Language—25.6%

Frequently asked questions

Is Magistral Small better than Phi-4?

Magistral Small and Phi-4 score almost the same on the Noometry Index (30.2 vs 31.2), so choose on price, context window or the category you care about most.

Which is cheaper, Magistral Small or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Magistral Small lists at $0.50 and $1.50.

Is Magistral Small or Phi-4 better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Magistral Small and Phi-4 share?

4 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper