Model comparison

Grok Build 0.1 vs MiniMax-M2

Grok Build 0.1 and MiniMax-M2 score almost the same on the Noometry Index (36.4 vs 37.4), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 19.4.
  • MiniMax-M2 is cheaper at $0.30 / $1.20 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 205K.
  • MiniMax-M2 has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and MiniMax-M2 specifications
Grok Build 0.1MiniMax-M2
ProviderxAIMiniMax
Noometry Index36.437.4
Released2026-04-162025-10-27
WeightsProprietaryOpen
Context window256K205K
Max output256K131K
Input $ / M tokens$1$0.30
Output $ / M tokens$2$1.20
Results tracked321

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), MiniMax-M2: 39.3 (#159)

Coding benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
SWE-bench Verified (bash only)—61%
LMArena WebDev—1297
SciCode50.2%—
LMArena Coding—1370

Agentic & Tool Use MiniMax-M2 leads

Grok Build 0.1: 22.7 (#129), MiniMax-M2: 25.1 (#109)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
Terminal-Bench—30%
GBAEval2.4%—
Vending-Bench 2—160.6

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), MiniMax-M2: 19.4 (#258)

Reasoning benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
Kagi LLM Benchmark—57.8%
NYT Connections (extended)—14.8%
CritPt9.1%—
LMArena Hard Prompts—1357

Math Not comparable

Grok Build 0.1: —, MiniMax-M2: 37.3 (#160)

Math benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
LMArena Math—1352

Knowledge Not comparable

Grok Build 0.1: —, MiniMax-M2: 37.0 (#163)

Knowledge benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
LMArena Expert—1337

Multilingual Not comparable

Grok Build 0.1: —, MiniMax-M2: 45.3 (#171)

Multilingual benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
LMArena Non-English—1313
LMArena Chinese—1366
LMArena French—1335
LMArena German—1355
LMArena Russian—1331
LMArena Spanish—1326

Instruction Following Not comparable

Grok Build 0.1: —, MiniMax-M2: 70.2 (#166)

Instruction Following benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
LMArena Instruction Following—1328

Long Context Not comparable

Grok Build 0.1: —, MiniMax-M2: 40.5 (#153)

Long Context benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
LMArena Longer Query—1331

Writing & Preference Not comparable

Grok Build 0.1: —, MiniMax-M2: 53.0 (#162)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1MiniMax-M2
LMArena Text—1340
LMArena Creative Writing—1286
LMArena Multi-Turn—1361

Frequently asked questions

Is Grok Build 0.1 better than MiniMax-M2?

Grok Build 0.1 and MiniMax-M2 score almost the same on the Noometry Index (36.4 vs 37.4), so choose on price, context window or the category you care about most.

Which is cheaper, Grok Build 0.1 or MiniMax-M2?

MiniMax-M2 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or MiniMax-M2 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 39.3 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 205K.

How many benchmarks do Grok Build 0.1 and MiniMax-M2 share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and MiniMax-M2 has 21.

Related comparisons

Go deeper