Model comparison

GPT-4.5 vs MiniMax-M2.7

GPT-4.5 and MiniMax-M2.7 score almost the same on the Noometry Index (37.2 vs 37.7), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • They share 18 benchmarks with published results for both. GPT-4.5 scores higher in 4 categories and MiniMax-M2.7 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-4.5 leads 32.6 to 25.9.
  • MiniMax-M2.7 has downloadable open weights; the other is API-only.

Side by side

GPT-4.5 and MiniMax-M2.7 specifications
GPT-4.5MiniMax-M2.7
ProviderOpenAIMiniMax
Noometry Index37.237.7
Released2025-02-272026-03-18
WeightsProprietaryOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked4230

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4.5: 42.2 (#109), MiniMax-M2.7: 41.8 (#120)

Coding benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
WeirdML39.4%37%
LMArena Coding13961454
Aider Polyglot44.9%—
LMArena WebDev—1398
SciCode—47%
LiveBench Coding75.2%—
ALE-Bench—599.25

Agentic & Tool Use GPT-4.5 leads

GPT-4.5: 27.9 (#97), MiniMax-M2.7: 25.1 (#111)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
Terminal-Bench—45.1%
Cybench17.5%—
ExploitBench—13.3%
GBAEval—0%

Reasoning MiniMax-M2.7 leads

GPT-4.5: 13.9 (#330), MiniMax-M2.7: 19.7 (#253)

Reasoning benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Hard Prompts14031422
Epoch Capabilities Index136.74145.85
ARC-AGI-20.8%—
SimpleBench34.5%—
NYT Connections (extended)—24.7%
ARC-AGI-110.3%—
CritPt—0.6%
EnigmaEval3.2%—
Thematic Generalization—39.3%
LiveBench Reasoning71.1%—
LiveBench Data Analysis64.3%—
ForecastBench61.7—
LiveBench69%—

Math GPT-4.5 leads

GPT-4.5: 32.6 (#211), MiniMax-M2.7: 25.9 (#263)

Math benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Math14121420
OTIS Mock AIME 2024-202537.8%—
ProofBench—3%
LiveBench Math69.3%—
MATH Level 578.6%—

Knowledge MiniMax-M2.7 leads

GPT-4.5: 32.5 (#211), MiniMax-M2.7: 37.7 (#152)

Knowledge benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Expert13941444
GPQA Diamond68.7%—
Humanity's Last Exam5.4%—
Confabulations13.6%—
Vectara Hallucination Rate—12.9%

Multimodal Not comparable

GPT-4.5: 37.6 (#71), MiniMax-M2.7: —

Multimodal benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Vision1195—
VPCT45%—

Multilingual GPT-4.5 leads

GPT-4.5: 52.5 (#83), MiniMax-M2.7: 50.3 (#123)

Multilingual benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Non-English14131382
LMArena Chinese14211441
LMArena French14181421
LMArena German14571398
LMArena Japanese14161262
LMArena Korean13921313
LMArena Russian14191383
LMArena Spanish—1403

Instruction Following MiniMax-M2.7 leads

GPT-4.5: 72.6 (#134), MiniMax-M2.7: 74.1 (#103)

Instruction Following benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Instruction Following14041405
LiveBench Instruction Following72.3%—

Long Context MiniMax-M2.7 leads

GPT-4.5: 40.4 (#155), MiniMax-M2.7: 43.3 (#99)

Long Context benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Longer Query14061419
Fiction.LiveBench63.9%—

Writing & Preference MiniMax-M2.7 leads

GPT-4.5: 56.9 (#134), MiniMax-M2.7: 58.9 (#112)

Writing & Preference benchmarks
BenchmarkGPT-4.5MiniMax-M2.7
LMArena Text14171405
LMArena Creative Writing13941354
LMArena Multi-Turn14441412
Short-Story Creative Writing75.6%—
EQ-Bench Creative Writing1258—
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than MiniMax-M2.7?

GPT-4.5 and MiniMax-M2.7 score almost the same on the Noometry Index (37.2 vs 37.7), so choose on price, context window or the category you care about most.

Is GPT-4.5 or MiniMax-M2.7 better for coding?

They score almost the same on coding (42.2 vs 41.8); test both on your own repository before choosing.

How many benchmarks do GPT-4.5 and MiniMax-M2.7 share?

18 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and MiniMax-M2.7 has 30.

Related comparisons

Go deeper