Model comparison
Magistral Small vs phi-3-medium 14B
Magistral Small and phi-3-medium 14B score almost the same on the Noometry Index (30.2 vs 29.7), so choose on price, context window or the category you care about most.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. Magistral Small scores higher in 2 categories and phi-3-medium 14B in 1 category; 3 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Magistral Small leads 30.9 to 9.1.
- The biggest single-benchmark swing is GPQA Diamond: 56.1% for Magistral Small and 27.6% for phi-3-medium 14B.
Side by side
| Magistral Small | phi-3-medium 14B | |
|---|---|---|
| Provider | Mistral AI | Microsoft |
| Noometry Index | 30.2 | 29.7 |
| Released | 2025-06-10 | 2024-04-23 |
| Weights | Open | Open |
| Context window | 128K | — |
| Max output | 40K | — |
| Input $ / M tokens | $0.50 | — |
| Output $ / M tokens | $1.50 | — |
| Results tracked | 10 | 13 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Magistral Small leads
Magistral Small: 38.4 (#176), phi-3-medium 14B: 36.8 (#201)
| Benchmark | Magistral Small | phi-3-medium 14B |
|---|---|---|
| SciCode | 35.2% | — |
| BigCodeBench Instruct | — | 37.6% |
| BigCodeBench Complete | — | 48.7% |
Reasoning Not comparable
Magistral Small: 6.8 (#350), phi-3-medium 14B: —
| Benchmark | Magistral Small | phi-3-medium 14B |
|---|---|---|
| Epoch Capabilities Index | 133.19 | 121.23 |
| ARC-AGI-2 | 0% | — |
| Kagi LLM Benchmark | 6.3% | — |
| ARC-AGI-1 | 5% | — |
| CritPt | 0.3% | — |
| Chess Puzzles | 3% | — |
| DTBench | 61.3% | — |
| Adversarial NLI | — | 55.8% |
| BIG-Bench Hard | — | 81.4% |
| HellaSwag | — | 82.4% |
| WinoGrande | — | 81.5% |
Math phi-3-medium 14B leads
Magistral Small: 26.2 (#261), phi-3-medium 14B: 27.3 (#250)
| Benchmark | Magistral Small | phi-3-medium 14B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 30% | — |
| MATH Level 5 | — | 17.6% |
Knowledge Magistral Small leads
Magistral Small: 30.9 (#223), phi-3-medium 14B: 9.1 (#306)
| Benchmark | Magistral Small | phi-3-medium 14B |
|---|---|---|
| GPQA Diamond | 56.1% | 27.6% |
| ARC (AI2) Challenge | — | 91.6% |
| MMLU | — | 78% |
| OpenBookQA | — | 87.4% |
| TriviaQA | — | 73.9% |
Frequently asked questions
Is Magistral Small better than phi-3-medium 14B?
Magistral Small and phi-3-medium 14B score almost the same on the Noometry Index (30.2 vs 29.7), so choose on price, context window or the category you care about most.
Is Magistral Small or phi-3-medium 14B better for coding?
Magistral Small scores higher on coding benchmarks: 38.4 versus 36.8 in the Noometry coding category.
How many benchmarks do Magistral Small and phi-3-medium 14B share?
2 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and phi-3-medium 14B has 13.