Model comparison

Magistral Medium vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 35.2 on the Noometry Index.

Last verified . 13 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Magistral Medium scores higher in 1 category and Mercury 2 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2 leads 23.8 to 8.6.
  • Mercury 2 is cheaper at $0.25 / $0.75 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 128K.
  • Magistral Medium has downloadable open weights; the other is API-only.

Side by side

Magistral Medium and Mercury 2 specifications
Magistral MediumMercury 2
ProviderMistral AIInception
Noometry Index35.239.1
Released2025-03-172026-02-20
WeightsOpenProprietary
Context window262K128K
Max output16K50K
Input $ / M tokens$2$0.25
Output $ / M tokens$5$0.75
Results tracked2217

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Magistral Medium: 39.1 (#161), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkMagistral MediumMercury 2
SciCode39.2%38.7%
LMArena Coding13191391
LMArena WebDev—1171
WeirdML—43.2%
ALE-Bench—785.58

Reasoning Mercury 2 leads

Magistral Medium: 8.6 (#348), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkMagistral MediumMercury 2
CritPt0.3%0.8%
LMArena Hard Prompts12671362
ARC-AGI-20%—
Kagi LLM Benchmark16.2%—
ARC-AGI-16.1%—

Math Not comparable

Magistral Medium: 35.1 (#189), Mercury 2: —

Math benchmarks
BenchmarkMagistral MediumMercury 2
LMArena Math1250—

Knowledge Mercury 2 leads

Magistral Medium: 33.5 (#202), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkMagistral MediumMercury 2
LMArena Expert12231358
Vectara Hallucination Rate—12.3%

Multilingual Mercury 2 leads

Magistral Medium: 39.6 (#224), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkMagistral MediumMercury 2
LMArena Non-English12321331
LMArena Chinese12271417
LMArena Russian12241304
LMArena French1267—
LMArena German1248—
LMArena Japanese1175—
LMArena Korean1125—
LMArena Spanish1271—

Instruction Following Mercury 2 leads

Magistral Medium: 66.0 (#211), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkMagistral MediumMercury 2
LMArena Instruction Following12541329

Long Context Mercury 2 leads

Magistral Medium: 39.3 (#183), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkMagistral MediumMercury 2
LMArena Longer Query12951330

Writing & Preference Mercury 2 leads

Magistral Medium: 46.3 (#219), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkMagistral MediumMercury 2
LMArena Text12551355
LMArena Creative Writing12451289
LMArena Multi-Turn12751358

Frequently asked questions

Is Magistral Medium better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 35.2 on the Noometry Index.

Which is cheaper, Magistral Medium or Mercury 2?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Magistral Medium lists at $2 and $5.

Is Magistral Medium or Mercury 2 better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 128K.

How many benchmarks do Magistral Medium and Mercury 2 share?

13 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper