Model comparison
Magistral Small vs Mistral Nemo
Magistral Small is the stronger model overall, scoring 30.2 to 26.4 on the Noometry Index. Mistral Nemo costs 5.0× less per token, which makes it the better buy when Magistral Small's lead doesn't matter for your workload.
Last verified . 3 shared benchmarks.
Summary
- They share 3 benchmarks with published results for both. Magistral Small scores higher in 2 categories and Mistral Nemo in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Magistral Small leads 30.9 to 12.3.
- The biggest single-benchmark swing is GPQA Diamond: 56.1% for Magistral Small and 29.9% for Mistral Nemo.
- Mistral Nemo is cheaper at $0.15 / $0.15 per million input/output tokens, against $0.50 / $1.50 for Magistral Small.
Side by side
| Magistral Small | Mistral Nemo | |
|---|---|---|
| Provider | Mistral AI | Mistral AI |
| Noometry Index | 30.2 | 26.4 |
| Released | 2025-06-10 | 2024-07-01 |
| Weights | Open | Open |
| Context window | 128K | 128K |
| Max output | 40K | 128K |
| Input $ / M tokens | $0.50 | $0.15 |
| Output $ / M tokens | $1.50 | $0.15 |
| Results tracked | 10 | 10 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Magistral Small: 38.4 (#176), Mistral Nemo: —
| Benchmark | Magistral Small | Mistral Nemo |
|---|---|---|
| SciCode | 35.2% | — |
Agentic & Tool Use Not comparable
Magistral Small: —, Mistral Nemo: 23.5 (#125)
| Benchmark | Magistral Small | Mistral Nemo |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 27.6% |
| BALROG | — | 17.6% |
Reasoning Mistral Nemo leads
Magistral Small: 6.8 (#350), Mistral Nemo: 20.7 (#232)
| Benchmark | Magistral Small | Mistral Nemo |
|---|---|---|
| DTBench | 61.3% | 48.6% |
| Epoch Capabilities Index | 133.19 | 118.68 |
| ARC-AGI-2 | 0% | — |
| Kagi LLM Benchmark | 6.3% | — |
| ARC-AGI-1 | 5% | — |
| CritPt | 0.3% | — |
| Chess Puzzles | 3% | — |
| PIQA | — | 83.5% |
Math Too close to call
Magistral Small: 26.2 (#261), Mistral Nemo: 25.5 (#268)
| Benchmark | Magistral Small | Mistral Nemo |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 30% | — |
| MATH Level 5 | — | 10.8% |
| GSM8K | — | 84.2% |
Knowledge Magistral Small leads
Magistral Small: 30.9 (#223), Mistral Nemo: 12.3 (#298)
| Benchmark | Magistral Small | Mistral Nemo |
|---|---|---|
| GPQA Diamond | 56.1% | 29.9% |
| BoolQ | — | 82.5% |
Writing & Preference Not comparable
Magistral Small: —, Mistral Nemo: 28.5 (#296)
| Benchmark | Magistral Small | Mistral Nemo |
|---|---|---|
| EQ-Bench Creative Writing | — | 881 |
Frequently asked questions
Is Magistral Small better than Mistral Nemo?
Magistral Small is the stronger model overall, scoring 30.2 to 26.4 on the Noometry Index. Mistral Nemo costs 5.0× less per token, which makes it the better buy when Magistral Small's lead doesn't matter for your workload.
Which is cheaper, Magistral Small or Mistral Nemo?
Mistral Nemo is cheaper. It lists at $0.15 per million input tokens and $0.15 per million output tokens; Magistral Small lists at $0.50 and $1.50.
Which has the bigger context window?
Both accept 128K tokens.
How many benchmarks do Magistral Small and Mistral Nemo share?
3 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Mistral Nemo has 10.