Model comparison

Mercury 2 vs Mistral Small

Mercury 2 is the stronger model overall, scoring 39.1 to 33.4 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mercury 2 scores higher in 6 categories and Mistral Small in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mercury 2 leads 36.2 to 31.0.
  • The biggest single-benchmark swing is SciCode: 38.7% for Mercury 2 and 26.5% for Mistral Small.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.25 / $0.75 for Mercury 2.
  • Mistral Small accepts more context: 262K tokens versus 128K.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Mercury 2 and Mistral Small specifications
Mercury 2Mistral Small
ProviderInceptionMistral AI
Noometry Index39.133.4
Released2026-02-202024-02-26
WeightsProprietaryOpen
Context window128K262K
Max output50K256K
Input $ / M tokens$0.25$0.15
Output $ / M tokens$0.75$0.60
Results tracked1739

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury 2: 33.5 (#255), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkMercury 2Mistral Small
SciCode38.7%26.5%
LMArena Coding13911362
ALE-Bench785.58497.62
LMArena WebDev1171—
WeirdML43.2%—
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%

Agentic & Tool Use Not comparable

Mercury 2: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkMercury 2Mistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Mercury 2 leads

Mercury 2: 23.8 (#170), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkMercury 2Mistral Small
CritPt0.8%0%
LMArena Hard Prompts13621335
Kagi LLM Benchmark—37.8%
LiveBench Reasoning—44.8%
DTBench—70.9%
LiveBench Data Analysis—53.7%
LMCA—20.6%
LiveBench—44%

Math Not comparable

Mercury 2: —, Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkMercury 2Mistral Small
OTIS Mock AIME 2024-2025—5.8%
LiveBench Math—39.9%
LMArena Math—1341
MATH Level 5—46.8%

Knowledge Mercury 2 leads

Mercury 2: 36.2 (#172), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkMercury 2Mistral Small
Vectara Hallucination Rate12.3%5.1%
LMArena Expert13581291
GPQA Diamond—47.5%
MMLU—68.7%

Multimodal Not comparable

Mercury 2: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkMercury 2Mistral Small
LMArena Vision—1142

Multilingual Mercury 2 leads

Mercury 2: 46.6 (#157), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkMercury 2Mistral Small
LMArena Non-English13311315
LMArena Chinese14171340
LMArena Russian13041324
LMArena French—1337
LMArena German—1340
LMArena Japanese—1275
LMArena Korean—1259
LMArena Spanish—1346

Instruction Following Mercury 2 leads

Mercury 2: 70.2 (#165), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkMercury 2Mistral Small
LMArena Instruction Following13291310
LiveBench Instruction Following—63.7%

Long Context Too close to call

Mercury 2: 40.5 (#154), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkMercury 2Mistral Small
LMArena Longer Query13301327

Writing & Preference Mercury 2 leads

Mercury 2: 53.8 (#155), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkMercury 2Mistral Small
LMArena Text13551338
LMArena Creative Writing12891305
LMArena Multi-Turn13581344
LiveBench Language—30.5%

Frequently asked questions

Is Mercury 2 better than Mistral Small?

Mercury 2 is the stronger model overall, scoring 39.1 to 33.4 on the Noometry Index.

Which is cheaper, Mercury 2 or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Mercury 2 lists at $0.25 and $0.75.

Is Mercury 2 or Mistral Small better for coding?

They score almost the same on coding (33.5 vs 34.0); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 128K.

How many benchmarks do Mercury 2 and Mistral Small share?

15 benchmarks have published results for both models. Mercury 2 has 17 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper