Model comparison

Mercury vs Mistral Medium

Mercury is the stronger model overall, scoring 37.6 to 36.3 on the Noometry Index.

Last verified . 9 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Mercury scores higher in 1 category and Mistral Medium in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 46.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 21.6% for Mercury and 50% for Mistral Medium.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mercury and Mistral Medium specifications
MercuryMistral Medium
ProviderInceptionMistral AI
Noometry Index37.636.3
Released—2023-12-11
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkMercuryMistral Medium
LMArena Coding13221434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Mercury: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkMercuryMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Mercury: 17.5 (#293), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkMercuryMistral Medium
Kagi LLM Benchmark21.6%50%
LMArena Hard Prompts12851426
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Not comparable

Mercury: —, Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkMercuryMistral Medium
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
LMArena Math—1408
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Mercury: —, Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkMercuryMistral Medium
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%
LMArena Expert—1408

Multimodal Not comparable

Mercury: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkMercuryMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Mercury: 41.6 (#206), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkMercuryMistral Medium
LMArena Non-English12601408
LMArena Chinese—1447
LMArena French—1459
LMArena German—1432
LMArena Japanese—1378
LMArena Korean—1380
LMArena Russian—1411
LMArena Spanish—1433

Instruction Following Mistral Medium leads

Mercury: 65.2 (#224), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkMercuryMistral Medium
LMArena Instruction Following12391398

Long Context Mistral Medium leads

Mercury: 38.4 (#198), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkMercuryMistral Medium
LMArena Longer Query12661406

Writing & Preference Mistral Medium leads

Mercury: 46.2 (#221), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkMercuryMistral Medium
LMArena Text12821424
LMArena Creative Writing11911391
LMArena Multi-Turn12821418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Mercury better than Mistral Medium?

Mercury is the stronger model overall, scoring 37.6 to 36.3 on the Noometry Index.

Is Mercury or Mistral Medium better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 34.2 in the Noometry coding category.

How many benchmarks do Mercury and Mistral Medium share?

9 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper