Model comparison

Claude 2.1 vs MiniMax-M2.7

MiniMax-M2.7 is the stronger model overall, scoring 37.7 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 1 category and MiniMax-M2.7 in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiniMax-M2.7 leads 37.7 to 15.4.
  • The biggest single-benchmark swing is WeirdML: 7.1% for Claude 2.1 and 37% for MiniMax-M2.7.
  • MiniMax-M2.7 has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and MiniMax-M2.7 specifications
Claude 2.1MiniMax-M2.7
ProviderAnthropicMiniMax
Noometry Index25.237.7
Released2023-11-212026-03-18
WeightsProprietaryOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked730

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.7 leads

Claude 2.1: 26.2 (#327), MiniMax-M2.7: 41.8 (#120)

Coding benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
WeirdML7.1%37%
LMArena WebDev—1398
SciCode—47%
LMArena Coding—1454
ALE-Bench—599.25

Agentic & Tool Use Not comparable

Claude 2.1: —, MiniMax-M2.7: 25.1 (#111)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
Terminal-Bench—45.1%
ExploitBench—13.3%
GBAEval—0%

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), MiniMax-M2.7: 19.7 (#253)

Reasoning benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
Epoch Capabilities Index119.27145.85
NYT Connections (extended)—24.7%
CritPt—0.6%
Thematic Generalization—39.3%
LMArena Hard Prompts—1422
DTBench51%—
ForecastBench54.2—

Math MiniMax-M2.7 leads

Claude 2.1: 10.2 (#315), MiniMax-M2.7: 25.9 (#263)

Math benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
OTIS Mock AIME 2024-20251.9%—
ProofBench—3%
LMArena Math—1420

Knowledge MiniMax-M2.7 leads

Claude 2.1: 15.4 (#292), MiniMax-M2.7: 37.7 (#152)

Knowledge benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
GPQA Diamond33%—
Vectara Hallucination Rate—12.9%
LMArena Expert—1444
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, MiniMax-M2.7: 50.3 (#123)

Multilingual benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
LMArena Non-English—1382
LMArena Chinese—1441
LMArena French—1421
LMArena German—1398
LMArena Japanese—1262
LMArena Korean—1313
LMArena Russian—1383
LMArena Spanish—1403

Instruction Following Not comparable

Claude 2.1: —, MiniMax-M2.7: 74.1 (#103)

Instruction Following benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
LMArena Instruction Following—1405

Long Context Not comparable

Claude 2.1: —, MiniMax-M2.7: 43.3 (#99)

Long Context benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
LMArena Longer Query—1419

Writing & Preference Not comparable

Claude 2.1: —, MiniMax-M2.7: 58.9 (#112)

Writing & Preference benchmarks
BenchmarkClaude 2.1MiniMax-M2.7
LMArena Text—1405
LMArena Creative Writing—1354
LMArena Multi-Turn—1412

Frequently asked questions

Is Claude 2.1 better than MiniMax-M2.7?

MiniMax-M2.7 is the stronger model overall, scoring 37.7 to 25.2 on the Noometry Index.

Is Claude 2.1 or MiniMax-M2.7 better for coding?

MiniMax-M2.7 scores higher on coding benchmarks: 41.8 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and MiniMax-M2.7 share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and MiniMax-M2.7 has 30.

Related comparisons

Go deeper