Model comparison

DeepSeek-V3.1 vs MiniMax-M2.5

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 38.3 on the Noometry Index.

Last verified . 21 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

MiniMax-M2.5 MiniMax

38.3

Rank #188 Confirmed

Summary

  • They share 21 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 6 categories and MiniMax-M2.5 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-V3.1 leads 38.9 to 26.9.
  • DeepSeek-V3.1 is cheaper at $0.25 / $0.95 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.5.
  • MiniMax-M2.5 accepts more context: 205K tokens versus 164K.

Side by side

DeepSeek-V3.1 and MiniMax-M2.5 specifications
DeepSeek-V3.1MiniMax-M2.5
ProviderDeepSeekMiniMax
Noometry Index42.838.3
Released2025-08-212026-02-12
WeightsOpenOpen
Context window164K205K
Max output8K131K
Input $ / M tokens$0.25$0.30
Output $ / M tokens$0.95$1.20
Results tracked2733

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.5 leads

DeepSeek-V3.1: 40.3 (#144), MiniMax-M2.5: 48.1 (#58)

Coding benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
LMArena Coding14171381
SWE-bench Verified (bash only)—75.8%
LMArena WebDev—1387
SWE-bench Multilingual—68.3%
WeirdML38.4%—
ALE-Bench—618.17

Agentic & Tool Use Not comparable

DeepSeek-V3.1: —, MiniMax-M2.5: 30.4 (#77)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
Terminal-Bench—42.7%
Vending-Bench 2—-23.16

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), MiniMax-M2.5: 17.5 (#292)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
Kagi LLM Benchmark53.2%55.2%
LMArena Hard Prompts14171372
Epoch Capabilities Index139.92146.68
ARC-AGI-2—4.9%
SimpleBench40%—
NYT Connections (extended)—16.8%
ARC-AGI-1—63.7%
DTBench82.7%—
LMCA24.3%—
ForecastBench58—

Math DeepSeek-V3.1 leads

DeepSeek-V3.1: 38.9 (#122), MiniMax-M2.5: 26.9 (#253)

Math benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
LMArena Math14201378
ProofBench—4%

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), MiniMax-M2.5: 39.2 (#135)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
Vectara Hallucination Rate5.5%9.1%
LMArena Expert14051379

Multilingual DeepSeek-V3.1 leads

DeepSeek-V3.1: 51.6 (#106), MiniMax-M2.5: 47.1 (#152)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
LMArena Non-English14001338
LMArena Chinese14691393
LMArena French14471362
LMArena German14111362
LMArena Japanese13781171
LMArena Korean13371232
LMArena Russian14051358
LMArena Spanish14311354

Instruction Following DeepSeek-V3.1 leads

DeepSeek-V3.1: 73.9 (#110), MiniMax-M2.5: 71.5 (#148)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
LMArena Instruction Following14001353

Long Context MiniMax-M2.5 leads

DeepSeek-V3.1: 36.3 (#232), MiniMax-M2.5: 37.5 (#216)

Long Context benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
LMArena Longer Query14221366
Fiction.LiveBench52.8%—
CL-bench—11.4%
CL-bench Life—6.3%

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), MiniMax-M2.5: 53.9 (#153)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1MiniMax-M2.5
LMArena Text14201359
LMArena Creative Writing14011331
EQ-Bench Creative Writing14361361
LMArena Multi-Turn14081364

Frequently asked questions

Is DeepSeek-V3.1 better than MiniMax-M2.5?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 38.3 on the Noometry Index.

Which is cheaper, DeepSeek-V3.1 or MiniMax-M2.5?

DeepSeek-V3.1 is cheaper. It lists at $0.25 per million input tokens and $0.95 per million output tokens; MiniMax-M2.5 lists at $0.30 and $1.20.

Is DeepSeek-V3.1 or MiniMax-M2.5 better for coding?

MiniMax-M2.5 scores higher on coding benchmarks: 48.1 versus 40.3 in the Noometry coding category.

Which has the bigger context window?

MiniMax-M2.5 does, with 205K tokens against 164K.

How many benchmarks do DeepSeek-V3.1 and MiniMax-M2.5 share?

21 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and MiniMax-M2.5 has 33.

Related comparisons

Go deeper