Model comparison

Grok Build 0.1 vs MiniMax-M2.7

MiniMax-M2.7 is the stronger model overall, scoring 37.7 to 36.4 on the Noometry Index.

Last verified . 3 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Grok Build 0.1 scores higher in 2 categories and MiniMax-M2.7 in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 19.7.
  • The biggest single-benchmark swing is CritPt: 9.1% for Grok Build 0.1 and 0.6% for MiniMax-M2.7.
  • MiniMax-M2.7 is cheaper at $0.30 / $1.20 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 205K.
  • MiniMax-M2.7 has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and MiniMax-M2.7 specifications
Grok Build 0.1MiniMax-M2.7
ProviderxAIMiniMax
Noometry Index36.437.7
Released2026-04-162026-03-18
WeightsProprietaryOpen
Context window256K205K
Max output256K131K
Input $ / M tokens$1$0.30
Output $ / M tokens$2$1.20
Results tracked330

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), MiniMax-M2.7: 41.8 (#120)

Coding benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
SciCode50.2%47%
LMArena WebDev—1398
WeirdML—37%
LMArena Coding—1454
ALE-Bench—599.25

Agentic & Tool Use MiniMax-M2.7 leads

Grok Build 0.1: 22.7 (#129), MiniMax-M2.7: 25.1 (#111)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
GBAEval2.4%0%
Terminal-Bench—45.1%
ExploitBench—13.3%

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), MiniMax-M2.7: 19.7 (#253)

Reasoning benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
CritPt9.1%0.6%
NYT Connections (extended)—24.7%
Thematic Generalization—39.3%
LMArena Hard Prompts—1422
Epoch Capabilities Index—145.85

Math Not comparable

Grok Build 0.1: —, MiniMax-M2.7: 25.9 (#263)

Math benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
ProofBench—3%
LMArena Math—1420

Knowledge Not comparable

Grok Build 0.1: —, MiniMax-M2.7: 37.7 (#152)

Knowledge benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
Vectara Hallucination Rate—12.9%
LMArena Expert—1444

Multilingual Not comparable

Grok Build 0.1: —, MiniMax-M2.7: 50.3 (#123)

Multilingual benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
LMArena Non-English—1382
LMArena Chinese—1441
LMArena French—1421
LMArena German—1398
LMArena Japanese—1262
LMArena Korean—1313
LMArena Russian—1383
LMArena Spanish—1403

Instruction Following Not comparable

Grok Build 0.1: —, MiniMax-M2.7: 74.1 (#103)

Instruction Following benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
LMArena Instruction Following—1405

Long Context Not comparable

Grok Build 0.1: —, MiniMax-M2.7: 43.3 (#99)

Long Context benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
LMArena Longer Query—1419

Writing & Preference Not comparable

Grok Build 0.1: —, MiniMax-M2.7: 58.9 (#112)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1MiniMax-M2.7
LMArena Text—1405
LMArena Creative Writing—1354
LMArena Multi-Turn—1412

Frequently asked questions

Is Grok Build 0.1 better than MiniMax-M2.7?

MiniMax-M2.7 is the stronger model overall, scoring 37.7 to 36.4 on the Noometry Index.

Which is cheaper, Grok Build 0.1 or MiniMax-M2.7?

MiniMax-M2.7 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or MiniMax-M2.7 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 41.8 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 205K.

How many benchmarks do Grok Build 0.1 and MiniMax-M2.7 share?

3 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and MiniMax-M2.7 has 30.

Related comparisons

Go deeper