Model comparison

MiniMax-M2.7 vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 37.7 on the Noometry Index.

Last verified . 20 shared benchmarks.

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 20 benchmarks with published results for both. MiniMax-M2.7 scores higher in 0 categories and Muse Spark in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 37.7.
  • The biggest single-benchmark swing is ProofBench: 3% for MiniMax-M2.7 and 17% for Muse Spark.
  • MiniMax-M2.7 has downloadable open weights; the other is API-only.

Side by side

MiniMax-M2.7 and Muse Spark specifications
MiniMax-M2.7Muse Spark
ProviderMiniMaxMeta
Noometry Index37.750.6
Released2026-03-182026-04-08
WeightsOpenProprietary
Context window205K—
Max output131K—
Input $ / M tokens$0.30—
Output $ / M tokens$1.20—
Results tracked3027

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

MiniMax-M2.7: 41.8 (#120), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkMiniMax-M2.7Muse Spark
SciCode47%51.5%
LMArena Coding14541481
LMArena WebDev1398—
WeirdML37%—
ALE-Bench599.25—

Agentic & Tool Use Not comparable

MiniMax-M2.7: 25.1 (#111), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.7Muse Spark
Terminal-Bench45.1%—
ExploitBench13.3%—
GBAEval0%—

Reasoning Muse Spark leads

MiniMax-M2.7: 19.7 (#253), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkMiniMax-M2.7Muse Spark
CritPt0.6%11.3%
LMArena Hard Prompts14221474
Epoch Capabilities Index145.85152.04
NYT Connections (extended)24.7%—
Thematic Generalization39.3%—

Math Muse Spark leads

MiniMax-M2.7: 25.9 (#263), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkMiniMax-M2.7Muse Spark
ProofBench3%17%
LMArena Math14201455
OTIS Mock AIME 2024-2025—88.9%
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

MiniMax-M2.7: 37.7 (#152), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkMiniMax-M2.7Muse Spark
LMArena Expert14441457
GPQA Diamond—89.8%
Humanity's Last Exam—40.6%
Vectara Hallucination Rate12.9%—

Multimodal Not comparable

MiniMax-M2.7: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkMiniMax-M2.7Muse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Muse Spark leads

MiniMax-M2.7: 50.3 (#123), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkMiniMax-M2.7Muse Spark
LMArena Non-English13821464
LMArena Chinese14411509
LMArena French14211497
LMArena German13981497
LMArena Korean13131459
LMArena Russian13831466
LMArena Spanish14031472
LMArena Japanese1262—

Instruction Following Muse Spark leads

MiniMax-M2.7: 74.1 (#103), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkMiniMax-M2.7Muse Spark
LMArena Instruction Following14051442

Long Context Muse Spark leads

MiniMax-M2.7: 43.3 (#99), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkMiniMax-M2.7Muse Spark
LMArena Longer Query14191451

Writing & Preference Muse Spark leads

MiniMax-M2.7: 58.9 (#112), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.7Muse Spark
LMArena Text14051474
LMArena Creative Writing13541459
LMArena Multi-Turn14121477

Frequently asked questions

Is MiniMax-M2.7 better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 37.7 on the Noometry Index.

Is MiniMax-M2.7 or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 41.8 in the Noometry coding category.

How many benchmarks do MiniMax-M2.7 and Muse Spark share?

20 benchmarks have published results for both models. MiniMax-M2.7 has 30 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper