Model comparison

Magistral Small vs phi-3-medium 14B

Magistral Small and phi-3-medium 14B score almost the same on the Noometry Index (30.2 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. Magistral Small scores higher in 2 categories and phi-3-medium 14B in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Magistral Small leads 30.9 to 9.1.
  • The biggest single-benchmark swing is GPQA Diamond: 56.1% for Magistral Small and 27.6% for phi-3-medium 14B.

Side by side

Magistral Small and phi-3-medium 14B specifications
Magistral Smallphi-3-medium 14B
ProviderMistral AIMicrosoft
Noometry Index30.229.7
Released2025-06-102024-04-23
WeightsOpenOpen
Context window128K—
Max output40K—
Input $ / M tokens$0.50—
Output $ / M tokens$1.50—
Results tracked1013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Magistral Small: 38.4 (#176), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkMagistral Smallphi-3-medium 14B
SciCode35.2%—
BigCodeBench Instruct—37.6%
BigCodeBench Complete—48.7%

Reasoning Not comparable

Magistral Small: 6.8 (#350), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkMagistral Smallphi-3-medium 14B
Epoch Capabilities Index133.19121.23
ARC-AGI-20%—
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
CritPt0.3%—
Chess Puzzles3%—
DTBench61.3%—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
HellaSwag—82.4%
WinoGrande—81.5%

Math phi-3-medium 14B leads

Magistral Small: 26.2 (#261), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkMagistral Smallphi-3-medium 14B
OTIS Mock AIME 2024-202530%—
MATH Level 5—17.6%

Knowledge Magistral Small leads

Magistral Small: 30.9 (#223), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkMagistral Smallphi-3-medium 14B
GPQA Diamond56.1%27.6%
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Frequently asked questions

Is Magistral Small better than phi-3-medium 14B?

Magistral Small and phi-3-medium 14B score almost the same on the Noometry Index (30.2 vs 29.7), so choose on price, context window or the category you care about most.

Is Magistral Small or phi-3-medium 14B better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 36.8 in the Noometry coding category.

How many benchmarks do Magistral Small and phi-3-medium 14B share?

2 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper