Model comparison

Mercury 2 vs MiniMax-M2.7

Mercury 2 is the stronger model overall, scoring 39.1 to 37.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mercury 2 scores higher in 1 category and MiniMax-M2.7 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where MiniMax-M2.7 leads 41.8 to 33.5.
  • The biggest single-benchmark swing is SciCode: 38.7% for Mercury 2 and 47% for MiniMax-M2.7.
  • Mercury 2 is cheaper at $0.25 / $0.75 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.7.
  • MiniMax-M2.7 accepts more context: 205K tokens versus 128K.
  • MiniMax-M2.7 has downloadable open weights; the other is API-only.

Side by side

Mercury 2 and MiniMax-M2.7 specifications
Mercury 2MiniMax-M2.7
ProviderInceptionMiniMax
Noometry Index39.137.7
Released2026-02-202026-03-18
WeightsProprietaryOpen
Context window128K205K
Max output50K131K
Input $ / M tokens$0.25$0.30
Output $ / M tokens$0.75$1.20
Results tracked1730

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.7 leads

Mercury 2: 33.5 (#255), MiniMax-M2.7: 41.8 (#120)

Coding benchmarks
BenchmarkMercury 2MiniMax-M2.7
LMArena WebDev11711398
SciCode38.7%47%
WeirdML43.2%37%
LMArena Coding13911454
ALE-Bench785.58599.25

Agentic & Tool Use Not comparable

Mercury 2: —, MiniMax-M2.7: 25.1 (#111)

Agentic & Tool Use benchmarks
BenchmarkMercury 2MiniMax-M2.7
Terminal-Bench—45.1%
ExploitBench—13.3%
GBAEval—0%

Reasoning Mercury 2 leads

Mercury 2: 23.8 (#170), MiniMax-M2.7: 19.7 (#253)

Reasoning benchmarks
BenchmarkMercury 2MiniMax-M2.7
CritPt0.8%0.6%
LMArena Hard Prompts13621422
NYT Connections (extended)—24.7%
Thematic Generalization—39.3%
Epoch Capabilities Index—145.85

Math Not comparable

Mercury 2: —, MiniMax-M2.7: 25.9 (#263)

Math benchmarks
BenchmarkMercury 2MiniMax-M2.7
ProofBench—3%
LMArena Math—1420

Knowledge MiniMax-M2.7 leads

Mercury 2: 36.2 (#172), MiniMax-M2.7: 37.7 (#152)

Knowledge benchmarks
BenchmarkMercury 2MiniMax-M2.7
Vectara Hallucination Rate12.3%12.9%
LMArena Expert13581444

Multilingual MiniMax-M2.7 leads

Mercury 2: 46.6 (#157), MiniMax-M2.7: 50.3 (#123)

Multilingual benchmarks
BenchmarkMercury 2MiniMax-M2.7
LMArena Non-English13311382
LMArena Chinese14171441
LMArena Russian13041383
LMArena French—1421
LMArena German—1398
LMArena Japanese—1262
LMArena Korean—1313
LMArena Spanish—1403

Instruction Following MiniMax-M2.7 leads

Mercury 2: 70.2 (#165), MiniMax-M2.7: 74.1 (#103)

Instruction Following benchmarks
BenchmarkMercury 2MiniMax-M2.7
LMArena Instruction Following13291405

Long Context MiniMax-M2.7 leads

Mercury 2: 40.5 (#154), MiniMax-M2.7: 43.3 (#99)

Long Context benchmarks
BenchmarkMercury 2MiniMax-M2.7
LMArena Longer Query13301419

Writing & Preference MiniMax-M2.7 leads

Mercury 2: 53.8 (#155), MiniMax-M2.7: 58.9 (#112)

Writing & Preference benchmarks
BenchmarkMercury 2MiniMax-M2.7
LMArena Text13551405
LMArena Creative Writing12891354
LMArena Multi-Turn13581412

Frequently asked questions

Is Mercury 2 better than MiniMax-M2.7?

Mercury 2 is the stronger model overall, scoring 39.1 to 37.7 on the Noometry Index.

Which is cheaper, Mercury 2 or MiniMax-M2.7?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; MiniMax-M2.7 lists at $0.30 and $1.20.

Is Mercury 2 or MiniMax-M2.7 better for coding?

MiniMax-M2.7 scores higher on coding benchmarks: 41.8 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

MiniMax-M2.7 does, with 205K tokens against 128K.

How many benchmarks do Mercury 2 and MiniMax-M2.7 share?

17 benchmarks have published results for both models. Mercury 2 has 17 scored results on Noometry and MiniMax-M2.7 has 30.

Related comparisons

Go deeper