Model comparison

Mercury 2.5 vs Mistral Medium 3.1

Mercury 2.5 is the stronger model overall, scoring 33.5 to 31.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • The widest gap is in reasoning, where Mercury 2.5 leads 22.4 to 10.6.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.40 / $2 for Mistral Medium 3.1.
  • Mercury 2.5 accepts more context: 260K tokens versus 131K.

Side by side

Mercury 2.5 and Mistral Medium 3.1 specifications
Mercury 2.5Mistral Medium 3.1
ProviderInceptionMistral AI
Noometry Index33.531.9
Released2026-09-08—
WeightsProprietaryProprietary
Context window260K131K
Max output66K105K
Input $ / M tokens$0.04$0.40
Output $ / M tokens$0.15$2
Results tracked43

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mercury 2.5: 39.5 (#156), Mistral Medium 3.1: —

Coding benchmarks
BenchmarkMercury 2.5Mistral Medium 3.1
SciCode38.5%—
ALE-Bench301.65—

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
BenchmarkMercury 2.5Mistral Medium 3.1
NYT Connections (extended)—6.5%
CritPt0%—
Thematic Generalization—20.3%

Math Not comparable

Mercury 2.5: 23.3 (#272), Mistral Medium 3.1: —

Math benchmarks
BenchmarkMercury 2.5Mistral Medium 3.1
ProofBench3%—

Writing & Preference Not comparable

Mercury 2.5: —, Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
BenchmarkMercury 2.5Mistral Medium 3.1
EQ-Bench Creative Writing—1476

Frequently asked questions

Is Mercury 2.5 better than Mistral Medium 3.1?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 31.9 on the Noometry Index.

Which is cheaper, Mercury 2.5 or Mistral Medium 3.1?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Mistral Medium 3.1 lists at $0.40 and $2.

Which has the bigger context window?

Mercury 2.5 does, with 260K tokens against 131K.

How many benchmarks do Mercury 2.5 and Mistral Medium 3.1 share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper