Model comparison

Mercury 2.5 vs MiniMax-M2

MiniMax-M2 is the stronger model overall, scoring 37.4 to 33.5 on the Noometry Index. Mercury 2.5 costs 7.8× less per token, which makes it the better buy when MiniMax-M2's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Summary

  • The widest gap is in math, where MiniMax-M2 leads 37.3 to 23.3.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.
  • Mercury 2.5 accepts more context: 260K tokens versus 205K.
  • MiniMax-M2 has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and MiniMax-M2 specifications
Mercury 2.5MiniMax-M2
ProviderInceptionMiniMax
Noometry Index33.537.4
Released2026-09-082025-10-27
WeightsProprietaryOpen
Context window260K205K
Max output66K131K
Input $ / M tokens$0.04$0.30
Output $ / M tokens$0.15$1.20
Results tracked421

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury 2.5: 39.5 (#156), MiniMax-M2: 39.3 (#159)

Coding benchmarks
BenchmarkMercury 2.5MiniMax-M2
SWE-bench Verified (bash only)—61%
LMArena WebDev—1297
SciCode38.5%—
LMArena Coding—1370
ALE-Bench301.65—

Agentic & Tool Use Not comparable

Mercury 2.5: —, MiniMax-M2: 25.1 (#109)

Agentic & Tool Use benchmarks
BenchmarkMercury 2.5MiniMax-M2
Terminal-Bench—30%
Vending-Bench 2—160.6

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), MiniMax-M2: 19.4 (#258)

Reasoning benchmarks
BenchmarkMercury 2.5MiniMax-M2
Kagi LLM Benchmark—57.8%
NYT Connections (extended)—14.8%
CritPt0%—
LMArena Hard Prompts—1357

Math MiniMax-M2 leads

Mercury 2.5: 23.3 (#272), MiniMax-M2: 37.3 (#160)

Math benchmarks
BenchmarkMercury 2.5MiniMax-M2
ProofBench3%—
LMArena Math—1352

Knowledge Not comparable

Mercury 2.5: —, MiniMax-M2: 37.0 (#163)

Knowledge benchmarks
BenchmarkMercury 2.5MiniMax-M2
LMArena Expert—1337

Multilingual Not comparable

Mercury 2.5: —, MiniMax-M2: 45.3 (#171)

Multilingual benchmarks
BenchmarkMercury 2.5MiniMax-M2
LMArena Non-English—1313
LMArena Chinese—1366
LMArena French—1335
LMArena German—1355
LMArena Russian—1331
LMArena Spanish—1326

Instruction Following Not comparable

Mercury 2.5: —, MiniMax-M2: 70.2 (#166)

Instruction Following benchmarks
BenchmarkMercury 2.5MiniMax-M2
LMArena Instruction Following—1328

Long Context Not comparable

Mercury 2.5: —, MiniMax-M2: 40.5 (#153)

Long Context benchmarks
BenchmarkMercury 2.5MiniMax-M2
LMArena Longer Query—1331

Writing & Preference Not comparable

Mercury 2.5: —, MiniMax-M2: 53.0 (#162)

Writing & Preference benchmarks
BenchmarkMercury 2.5MiniMax-M2
LMArena Text—1340
LMArena Creative Writing—1286
LMArena Multi-Turn—1361

Frequently asked questions

Is Mercury 2.5 better than MiniMax-M2?

MiniMax-M2 is the stronger model overall, scoring 37.4 to 33.5 on the Noometry Index. Mercury 2.5 costs 7.8× less per token, which makes it the better buy when MiniMax-M2's lead doesn't matter for your workload.

Which is cheaper, Mercury 2.5 or MiniMax-M2?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; MiniMax-M2 lists at $0.30 and $1.20.

Is Mercury 2.5 or MiniMax-M2 better for coding?

They score almost the same on coding (39.5 vs 39.3); test both on your own repository before choosing.

Which has the bigger context window?

Mercury 2.5 does, with 260K tokens against 205K.

How many benchmarks do Mercury 2.5 and MiniMax-M2 share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and MiniMax-M2 has 21.

Related comparisons

Go deeper