Model comparison

MiniMax-M2 vs Mistral Small 3.1

MiniMax-M2 is the stronger model overall, scoring 37.4 to 31.7 on the Noometry Index.

Last verified . 15 shared benchmarks.

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 15 benchmarks with published results for both. MiniMax-M2 scores higher in 7 categories and Mistral Small 3.1 in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where MiniMax-M2 leads 37.3 to 14.7.
  • Mistral Small 3.1 is cheaper at $0.35 / $0.56 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.
  • MiniMax-M2 accepts more context: 205K tokens versus 128K.

Side by side

MiniMax-M2 and Mistral Small 3.1 specifications
MiniMax-M2Mistral Small 3.1
ProviderMiniMaxMistral AI
Noometry Index37.431.7
Released2025-10-272025-03-17
WeightsOpenOpen
Context window205K128K
Max output131K102K
Input $ / M tokens$0.30$0.35
Output $ / M tokens$1.20$0.56
Results tracked2128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

MiniMax-M2: 39.3 (#159), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Coding13701309
SWE-bench Verified (bash only)61%—
LMArena WebDev1297—

Agentic & Tool Use Not comparable

MiniMax-M2: 25.1 (#109), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
Terminal-Bench30%—
Vending-Bench 2160.6—

Reasoning Too close to call

MiniMax-M2: 19.4 (#258), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Hard Prompts13571278
Kagi LLM Benchmark57.8%—
NYT Connections (extended)14.8%—
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math MiniMax-M2 leads

MiniMax-M2: 37.3 (#160), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Math13521262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%

Knowledge MiniMax-M2 leads

MiniMax-M2: 37.0 (#163), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Expert13371257
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%

Multimodal Not comparable

MiniMax-M2: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Vision—1136

Multilingual MiniMax-M2 leads

MiniMax-M2: 45.3 (#171), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Non-English13131255
LMArena Chinese13661253
LMArena French13351273
LMArena German13551266
LMArena Russian13311263
LMArena Spanish13261283
LMArena Japanese—1208
LMArena Korean—1206

Instruction Following MiniMax-M2 leads

MiniMax-M2: 70.2 (#166), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Instruction Following13281264
IFEval—75%

Long Context MiniMax-M2 leads

MiniMax-M2: 40.5 (#153), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Longer Query13311299

Writing & Preference MiniMax-M2 leads

MiniMax-M2: 53.0 (#162), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkMiniMax-M2Mistral Small 3.1
LMArena Text13401277
LMArena Creative Writing12861253
LMArena Multi-Turn13611270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is MiniMax-M2 better than Mistral Small 3.1?

MiniMax-M2 is the stronger model overall, scoring 37.4 to 31.7 on the Noometry Index.

Which is cheaper, MiniMax-M2 or Mistral Small 3.1?

Mistral Small 3.1 is cheaper. It lists at $0.35 per million input tokens and $0.56 per million output tokens; MiniMax-M2 lists at $0.30 and $1.20.

Is MiniMax-M2 or Mistral Small 3.1 better for coding?

They score almost the same on coding (39.3 vs 38.3); test both on your own repository before choosing.

Which has the bigger context window?

MiniMax-M2 does, with 205K tokens against 128K.

How many benchmarks do MiniMax-M2 and Mistral Small 3.1 share?

15 benchmarks have published results for both models. MiniMax-M2 has 21 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper