Model comparison

DeepSeek-R1-Distill-Qwen-1.5B vs MiniMax-M2.7

MiniMax-M2.7 is the stronger model overall, scoring 37.7 to 26.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • The widest gap is in knowledge, where MiniMax-M2.7 leads 37.7 to 16.0.

Side by side

DeepSeek-R1-Distill-Qwen-1.5B and MiniMax-M2.7 specifications
DeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
ProviderDeepSeekMiniMax
Noometry Index26.137.7
Released2025-01-202026-03-18
WeightsOpenOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked530

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.7 leads

DeepSeek-R1-Distill-Qwen-1.5B: 21.8 (#336), MiniMax-M2.7: 41.8 (#120)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
LMArena WebDev—1398
SciCode—47%
WeirdML—37%
BigCodeBench Instruct7%—
LMArena Coding—1454
BigCodeBench Complete7.9%—
ALE-Bench—599.25

Agentic & Tool Use Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, MiniMax-M2.7: 25.1 (#111)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
Terminal-Bench—45.1%
ExploitBench—13.3%
GBAEval—0%

Reasoning Too close to call

DeepSeek-R1-Distill-Qwen-1.5B: 19.2 (#262), MiniMax-M2.7: 19.7 (#253)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
NYT Connections (extended)—24.7%
CritPt—0.6%
Chess Puzzles0%—
Thematic Generalization—39.3%
LMArena Hard Prompts—1422
Epoch Capabilities Index—145.85

Math MiniMax-M2.7 leads

DeepSeek-R1-Distill-Qwen-1.5B: 23.0 (#274), MiniMax-M2.7: 25.9 (#263)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
OTIS Mock AIME 2024-202521.4%—
ProofBench—3%
LMArena Math—1420

Knowledge MiniMax-M2.7 leads

DeepSeek-R1-Distill-Qwen-1.5B: 16.0 (#290), MiniMax-M2.7: 37.7 (#152)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
GPQA Diamond33.6%—
Vectara Hallucination Rate—12.9%
LMArena Expert—1444

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, MiniMax-M2.7: 50.3 (#123)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
LMArena Non-English—1382
LMArena Chinese—1441
LMArena French—1421
LMArena German—1398
LMArena Japanese—1262
LMArena Korean—1313
LMArena Russian—1383
LMArena Spanish—1403

Instruction Following Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, MiniMax-M2.7: 74.1 (#103)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
LMArena Instruction Following—1405

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, MiniMax-M2.7: 43.3 (#99)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
LMArena Longer Query—1419

Writing & Preference Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, MiniMax-M2.7: 58.9 (#112)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BMiniMax-M2.7
LMArena Text—1405
LMArena Creative Writing—1354
LMArena Multi-Turn—1412

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-1.5B better than MiniMax-M2.7?

MiniMax-M2.7 is the stronger model overall, scoring 37.7 to 26.1 on the Noometry Index.

Is DeepSeek-R1-Distill-Qwen-1.5B or MiniMax-M2.7 better for coding?

MiniMax-M2.7 scores higher on coding benchmarks: 41.8 versus 21.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Qwen-1.5B and MiniMax-M2.7 share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-1.5B has 5 scored results on Noometry and MiniMax-M2.7 has 30.

Related comparisons

Go deeper