Model comparison

GPT-5.1 vs MiniMax-M2.7

GPT-5.1 is the stronger model overall, scoring 49.0 to 37.7 on the Noometry Index. MiniMax-M2.7 costs 6.5× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Last verified . 25 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GPT-5.1 scores higher in 9 categories and MiniMax-M2.7 in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.1 leads 52.2 to 25.9.
  • The biggest single-benchmark swing is WeirdML: 60.8% for GPT-5.1 and 37% for MiniMax-M2.7.
  • MiniMax-M2.7 is cheaper at $0.30 / $1.20 per million input/output tokens, against $1.25 / $10 for GPT-5.1.
  • GPT-5.1 accepts more context: 400K tokens versus 205K.
  • MiniMax-M2.7 has downloadable open weights; the other is API-only.

Side by side

GPT-5.1 and MiniMax-M2.7 specifications
GPT-5.1MiniMax-M2.7
ProviderOpenAIMiniMax
Noometry Index49.037.7
Released2025-11-132026-03-18
WeightsProprietaryOpen
Context window400K205K
Max output128K131K
Input $ / M tokens$1.25$0.30
Output $ / M tokens$10$1.20
Results tracked6330

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1 leads

GPT-5.1: 46.4 (#66), MiniMax-M2.7: 41.8 (#120)

Coding benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena WebDev13951398
SciCode43.3%47%
WeirdML60.8%37%
LMArena Coding14541454
ALE-Bench1,192599.25
SWE-bench Verified68%—
SWE-bench Verified (bash only)66%—
GSO13.7%—
LiveBench Coding72.5%—

Agentic & Tool Use GPT-5.1 leads

GPT-5.1: 32.7 (#60), MiniMax-M2.7: 25.1 (#111)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
Terminal-Bench47.6%45.1%
DeepResearch Bench42.8%—
ExploitBench—13.3%
GBAEval—0%
LMArena Search1199—
Vending-Bench 21,473—

Reasoning GPT-5.1 leads

GPT-5.1: 39.8 (#58), MiniMax-M2.7: 19.7 (#253)

Reasoning benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
CritPt4.9%0.6%
LMArena Hard Prompts14571422
Epoch Capabilities Index149.64145.85
ARC-AGI-217.6%—
SimpleBench53.2%—
NYT Connections (extended)—24.7%
ARC-AGI-172.8%—
Chess Puzzles32%—
EnigmaEval11.2%—
Thematic Generalization—39.3%
LiveBench Reasoning95.8%—
Mystery Game Puzzles19%—
DTBench90.1%—
LiveBench Data Analysis72.1%—
LMCA43.9%—
ForecastBench58.1—
LiveBench78.8%—

Math GPT-5.1 leads

GPT-5.1: 52.2 (#51), MiniMax-M2.7: 25.9 (#263)

Math benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena Math14471420
OTIS Mock AIME 2024-202588.6%—
ProofBench—3%
Omni-MATH46.4%—
LiveBench Math94.5%—
FrontierMath (Feb 2025 set)31%—
FrontierMath Tier 4 (v1)12.5%—

Knowledge GPT-5.1 leads

GPT-5.1: 50.6 (#71), MiniMax-M2.7: 37.7 (#152)

Knowledge benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
Vectara Hallucination Rate10.9%12.9%
LMArena Expert14701444
GPQA Diamond87.6%—
Humanity's Last Exam23.7%—
SimpleQA Verified48%—
MMLU-Pro57.9%—
GPQA (HELM)44.2%—

Multimodal Not comparable

GPT-5.1: 44.8 (#19), MiniMax-M2.7: —

Multimodal benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena Vision1250—
VPCT58.7%—
LMArena Document1403—

Multilingual GPT-5.1 leads

GPT-5.1: 53.8 (#56), MiniMax-M2.7: 50.3 (#123)

Multilingual benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena Non-English14311382
LMArena Chinese14951441
LMArena French14501421
LMArena German14381398
LMArena Japanese14531262
LMArena Korean14011313
LMArena Russian14351383
LMArena Spanish14331403

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), MiniMax-M2.7: 74.1 (#103)

Instruction Following benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena Instruction Following14431405
LiveBench Instruction Following93.3%—
IFEval93.5%—

Long Context GPT-5.1 leads

GPT-5.1: 47.6 (#14), MiniMax-M2.7: 43.3 (#99)

Long Context benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena Longer Query14471419
CL-bench23.7%—
CL-bench Life17.3%—

Writing & Preference GPT-5.1 leads

GPT-5.1: 64.5 (#55), MiniMax-M2.7: 58.9 (#112)

Writing & Preference benchmarks
BenchmarkGPT-5.1MiniMax-M2.7
LMArena Text14431405
LMArena Creative Writing14271354
LMArena Multi-Turn14501412
WildBench86.3%—
LiveBench Language80.2%—

Frequently asked questions

Is GPT-5.1 better than MiniMax-M2.7?

GPT-5.1 is the stronger model overall, scoring 49.0 to 37.7 on the Noometry Index. MiniMax-M2.7 costs 6.5× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Which is cheaper, GPT-5.1 or MiniMax-M2.7?

MiniMax-M2.7 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; GPT-5.1 lists at $1.25 and $10.

Is GPT-5.1 or MiniMax-M2.7 better for coding?

GPT-5.1 scores higher on coding benchmarks: 46.4 versus 41.8 in the Noometry coding category.

Which has the bigger context window?

GPT-5.1 does, with 400K tokens against 205K.

How many benchmarks do GPT-5.1 and MiniMax-M2.7 share?

25 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and MiniMax-M2.7 has 30.

Related comparisons

Go deeper