Model comparison

Mercury vs Mistral Small 3.1

Mercury is the stronger model overall, scoring 37.6 to 31.7 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 4 categories and Mistral Small 3.1 in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury leads 46.2 to 37.0.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Mercury and Mistral Small 3.1 specifications
MercuryMistral Small 3.1
ProviderInceptionMistral AI
Noometry Index37.631.7
Released—2025-03-17
WeightsProprietaryOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked928

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury: 38.7 (#170), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Coding13221309

Reasoning Mistral Small 3.1 leads

Mercury: 17.5 (#293), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Hard Prompts12851278
Kagi LLM Benchmark21.6%—
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math Not comparable

Mercury: —, Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkMercuryMistral Small 3.1
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%
LMArena Math—1262

Knowledge Not comparable

Mercury: —, Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkMercuryMistral Small 3.1
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%
LMArena Expert—1257

Multimodal Not comparable

Mercury: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Vision—1136

Multilingual Too close to call

Mercury: 41.6 (#206), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Non-English12601255
LMArena Chinese—1253
LMArena French—1273
LMArena German—1266
LMArena Japanese—1208
LMArena Korean—1206
LMArena Russian—1263
LMArena Spanish—1283

Instruction Following Mercury leads

Mercury: 65.2 (#224), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Instruction Following12391264
IFEval—75%

Long Context Mistral Small 3.1 leads

Mercury: 38.4 (#198), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Longer Query12661299

Writing & Preference Mercury leads

Mercury: 46.2 (#221), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkMercuryMistral Small 3.1
LMArena Text12821277
LMArena Creative Writing11911253
LMArena Multi-Turn12821270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is Mercury better than Mistral Small 3.1?

Mercury is the stronger model overall, scoring 37.6 to 31.7 on the Noometry Index.

Is Mercury or Mistral Small 3.1 better for coding?

They score almost the same on coding (38.7 vs 38.3); test both on your own repository before choosing.

How many benchmarks do Mercury and Mistral Small 3.1 share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper