Model comparison

Claude 3.7 Sonnet vs MiniMax M1

Claude 3.7 Sonnet and MiniMax M1 score almost the same on the Noometry Index (39.5 vs 40.3), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

MiniMax M1 MiniMax

40.3

Rank #150 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 5 categories and MiniMax M1 in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Claude 3.7 Sonnet leads 50.3 to 41.4.
  • The biggest single-benchmark swing is Fiction.LiveBench: 83.3% for Claude 3.7 Sonnet and 69.4% for MiniMax M1.
  • MiniMax M1 has downloadable open weights; the other is API-only.

Side by side

Claude 3.7 Sonnet and MiniMax M1 specifications
Claude 3.7 SonnetMiniMax M1
ProviderAnthropicMiniMax
Noometry Index39.540.3
Released2025-02-242025-06-13
WeightsProprietaryOpen
Context window—1M
Max output—40K
Input $ / M tokens—$0.55
Output $ / M tokens—$2.20
Results tracked5818

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.7 Sonnet: 40.6 (#136), MiniMax M1: 39.9 (#153)

Coding benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Coding13611359
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), MiniMax M1: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning MiniMax M1 leads

Claude 3.7 Sonnet: 18.6 (#277), MiniMax M1: 26.9 (#126)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Hard Prompts13331339
ARC-AGI-20.9%—
SimpleBench46.4%—
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Too close to call

Claude 3.7 Sonnet: 37.5 (#153), MiniMax M1: 37.5 (#151)

Math benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Math13371361
OTIS Mock AIME 2024-202557.8%—
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), MiniMax M1: 36.4 (#170)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Expert13211317
GPQA Diamond79.7%—
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), MiniMax M1: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual MiniMax M1 leads

Claude 3.7 Sonnet: 44.1 (#179), MiniMax M1: 45.8 (#163)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Non-English12961319
LMArena Chinese12991360
LMArena French13031370
LMArena German13011350
LMArena Japanese12671217
LMArena Korean12491266
LMArena Russian13111329
LMArena Spanish12981353

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), MiniMax M1: 69.3 (#174)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Instruction Following13521312
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), MiniMax M1: 41.4 (#141)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
Fiction.LiveBench83.3%69.4%
LMArena Longer Query13731326

Writing & Preference Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 54.4 (#150), MiniMax M1: 53.1 (#161)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetMiniMax M1
LMArena Text13141343
LMArena Creative Writing13321298
LMArena Multi-Turn13391335
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than MiniMax M1?

Claude 3.7 Sonnet and MiniMax M1 score almost the same on the Noometry Index (39.5 vs 40.3), so choose on price, context window or the category you care about most.

Is Claude 3.7 Sonnet or MiniMax M1 better for coding?

They score almost the same on coding (40.6 vs 39.9); test both on your own repository before choosing.

How many benchmarks do Claude 3.7 Sonnet and MiniMax M1 share?

18 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and MiniMax M1 has 18.

Related comparisons

Go deeper