Model comparison

Magistral Small vs Mercury 2.5

Mercury 2.5 is the stronger model overall, scoring 33.5 to 30.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • They share 2 benchmarks with published results for both. Magistral Small scores higher in 1 category and Mercury 2.5 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2.5 leads 22.4 to 6.8.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.50 / $1.50 for Magistral Small.
  • Mercury 2.5 accepts more context: 260K tokens versus 128K.
  • Magistral Small has downloadable open weights; the other is API-only.

Side by side

Magistral Small and Mercury 2.5 specifications
Magistral SmallMercury 2.5
ProviderMistral AIInception
Noometry Index30.233.5
Released2025-06-102026-09-08
WeightsOpenProprietary
Context window128K260K
Max output40K66K
Input $ / M tokens$0.50$0.04
Output $ / M tokens$1.50$0.15
Results tracked104

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Magistral Small: 38.4 (#176), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkMagistral SmallMercury 2.5
SciCode35.2%38.5%
ALE-Bench—301.65

Reasoning Mercury 2.5 leads

Magistral Small: 6.8 (#350), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkMagistral SmallMercury 2.5
CritPt0.3%0%
ARC-AGI-20%—
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
Chess Puzzles3%—
DTBench61.3%—
Epoch Capabilities Index133.19—

Math Magistral Small leads

Magistral Small: 26.2 (#261), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkMagistral SmallMercury 2.5
OTIS Mock AIME 2024-202530%—
ProofBench—3%

Knowledge Not comparable

Magistral Small: 30.9 (#223), Mercury 2.5: —

Knowledge benchmarks
BenchmarkMagistral SmallMercury 2.5
GPQA Diamond56.1%—

Frequently asked questions

Is Magistral Small better than Mercury 2.5?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 30.2 on the Noometry Index.

Which is cheaper, Magistral Small or Mercury 2.5?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Magistral Small lists at $0.50 and $1.50.

Is Magistral Small or Mercury 2.5 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 38.4 in the Noometry coding category.

Which has the bigger context window?

Mercury 2.5 does, with 260K tokens against 128K.

How many benchmarks do Magistral Small and Mercury 2.5 share?

2 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper