Model comparison
Inkling vs Magistral Small
Inkling is the stronger model overall, scoring 44.1 to 30.2 on the Noometry Index. Magistral Small costs 3.4× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.
Last verified . 9 shared benchmarks.
Summary
- They share 9 benchmarks with published results for both. Inkling scores higher in 3 categories and Magistral Small in 1 category; 4 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Inkling leads 40.4 to 6.8.
- The biggest single-benchmark swing is ARC-AGI-1: 79.5% for Inkling and 5% for Magistral Small.
- Magistral Small is cheaper at $0.50 / $1.50 per million input/output tokens, against $1.87 / $4.68 for Inkling.
- Magistral Small accepts more context: 128K tokens versus 66K.
Side by side
| Inkling | Magistral Small | |
|---|---|---|
| Provider | Thinking Machines Lab | Mistral AI |
| Noometry Index | 44.1 | 30.2 |
| Released | 2026-07-15 | 2025-06-10 |
| Weights | Open | Open |
| Context window | 66K | 128K |
| Max output | 66K | 40K |
| Input $ / M tokens | $1.87 | $0.50 |
| Output $ / M tokens | $4.68 | $1.50 |
| Results tracked | 41 | 10 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Magistral Small leads
Inkling: 34.5 (#234), Magistral Small: 38.4 (#176)
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| SciCode | 47% | 35.2% |
| FrontierCode | 14% | — |
| LMArena WebDev | 1413 | — |
| FrontierSWE | 4.1% | — |
| WeirdML | 32.3% | — |
| LMArena Coding | 1464 | — |
| ALE-Bench | 946 | — |
Agentic & Tool Use Not comparable
Inkling: 29.6 (#85), Magistral Small: —
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| APEX-Agents | 33.8% | — |
| τ²-bench Banking | 25% | — |
Reasoning Inkling leads
Inkling: 40.4 (#56), Magistral Small: 6.8 (#350)
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| ARC-AGI-2 | 36.5% | 0% |
| ARC-AGI-1 | 79.5% | 5% |
| CritPt | 5.4% | 0.3% |
| Chess Puzzles | 21% | 3% |
| DTBench | 87.5% | 61.3% |
| Epoch Capabilities Index | 148.54 | 133.19 |
| SimpleBench | 50% | — |
| Kagi LLM Benchmark | — | 6.3% |
| LMArena Hard Prompts | 1451 | — |
| LMCA | 37.6% | — |
Math Inkling leads
Inkling: 31.3 (#225), Magistral Small: 26.2 (#261)
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 88.9% | 30% |
| FrontierMath (Tiers 1-3) | 33.3% | — |
| FrontierMath Tier 4 | 4.9% | — |
| ProofBench | 0% | — |
| LMArena Math | 1479 | — |
Knowledge Inkling leads
Inkling: 55.1 (#49), Magistral Small: 30.9 (#223)
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| GPQA Diamond | 88.3% | 56.1% |
| SimpleQA Verified | 40.3% | — |
| LMArena Expert | 1465 | — |
Multilingual Not comparable
Inkling: 54.0 (#52), Magistral Small: —
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| LMArena Non-English | 1434 | — |
| LMArena Chinese | 1490 | — |
| LMArena French | 1458 | — |
| LMArena German | 1446 | — |
| LMArena Japanese | 1429 | — |
| LMArena Korean | 1404 | — |
| LMArena Russian | 1429 | — |
| LMArena Spanish | 1448 | — |
Instruction Following Not comparable
Inkling: 75.1 (#71), Magistral Small: —
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| LMArena Instruction Following | 1426 | — |
Long Context Not comparable
Inkling: 43.8 (#86), Magistral Small: —
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| LMArena Longer Query | 1434 | — |
Writing & Preference Not comparable
Inkling: 65.2 (#51), Magistral Small: —
| Benchmark | Inkling | Magistral Small |
|---|---|---|
| LMArena Text | 1441 | — |
| LMArena Creative Writing | 1387 | — |
| EQ-Bench Creative Writing | 1611 | — |
| EQ-Bench 4 | 1226 | — |
| LMArena Multi-Turn | 1436 | — |
Frequently asked questions
Is Inkling better than Magistral Small?
Inkling is the stronger model overall, scoring 44.1 to 30.2 on the Noometry Index. Magistral Small costs 3.4× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.
Which is cheaper, Inkling or Magistral Small?
Magistral Small is cheaper. It lists at $0.50 per million input tokens and $1.50 per million output tokens; Inkling lists at $1.87 and $4.68.
Is Inkling or Magistral Small better for coding?
Magistral Small scores higher on coding benchmarks: 38.4 versus 34.5 in the Noometry coding category.
Which has the bigger context window?
Magistral Small does, with 128K tokens against 66K.
How many benchmarks do Inkling and Magistral Small share?
9 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Magistral Small has 10.