Model comparison

Grok-2 (Dec 2024) vs MiniMax M1

MiniMax M1 is the stronger model overall, scoring 40.3 to 33.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

MiniMax M1 MiniMax

40.3

Rank #150 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok-2 (Dec 2024) scores higher in 0 categories and MiniMax M1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where MiniMax M1 leads 37.5 to 20.8.
  • MiniMax M1 has downloadable open weights; the other is API-only.

Side by side

Grok-2 (Dec 2024) and MiniMax M1 specifications
Grok-2 (Dec 2024)MiniMax M1
ProviderxAIMiniMax
Noometry Index33.740.3
Released2024-08-132025-06-13
WeightsProprietaryOpen
Context window—1M
Max output—40K
Input $ / M tokens—$0.55
Output $ / M tokens—$2.20
Results tracked3418

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax M1 leads

Grok-2 (Dec 2024): 33.3 (#258), MiniMax M1: 39.9 (#153)

Coding benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Coding12871359
WeirdML22.2%—
LiveBench Coding46.4%—

Reasoning MiniMax M1 leads

Grok-2 (Dec 2024): 16.9 (#299), MiniMax M1: 26.9 (#126)

Reasoning benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Hard Prompts12721339
SimpleBench22.7%—
LiveBench Reasoning54.8%—
DTBench65.2%—
LiveBench Data Analysis54.5%—
Epoch Capabilities Index130.48—
LiveBench54.3%—

Math MiniMax M1 leads

Grok-2 (Dec 2024): 20.8 (#284), MiniMax M1: 37.5 (#151)

Math benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Math12831361
OTIS Mock AIME 2024-202511.5%—
LiveBench Math54.9%—
MATH Level 563.5%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge MiniMax M1 leads

Grok-2 (Dec 2024): 29.8 (#233), MiniMax M1: 36.4 (#170)

Knowledge benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Expert12541317
GPQA Diamond53.8%—
Confabulations20.1%—

Multilingual MiniMax M1 leads

Grok-2 (Dec 2024): 43.1 (#188), MiniMax M1: 45.8 (#163)

Multilingual benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Non-English12821319
LMArena Chinese12891360
LMArena French13181370
LMArena German12871350
LMArena Japanese12441217
LMArena Korean12371266
LMArena Russian12861329
LMArena Spanish12811353

Instruction Following MiniMax M1 leads

Grok-2 (Dec 2024): 66.9 (#202), MiniMax M1: 69.3 (#174)

Instruction Following benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Instruction Following12701312
LiveBench Instruction Following69.6%—

Long Context MiniMax M1 leads

Grok-2 (Dec 2024): 38.8 (#190), MiniMax M1: 41.4 (#141)

Long Context benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Longer Query12761326
Fiction.LiveBench—69.4%

Writing & Preference MiniMax M1 leads

Grok-2 (Dec 2024): 48.6 (#198), MiniMax M1: 53.1 (#161)

Writing & Preference benchmarks
BenchmarkGrok-2 (Dec 2024)MiniMax M1
LMArena Text13051343
LMArena Creative Writing12841298
LMArena Multi-Turn12901335
Short-Story Creative Writing63.6%—
LiveBench Language45.6%—

Frequently asked questions

Is Grok-2 (Dec 2024) better than MiniMax M1?

MiniMax M1 is the stronger model overall, scoring 40.3 to 33.7 on the Noometry Index.

Is Grok-2 (Dec 2024) or MiniMax M1 better for coding?

MiniMax M1 scores higher on coding benchmarks: 39.9 versus 33.3 in the Noometry coding category.

How many benchmarks do Grok-2 (Dec 2024) and MiniMax M1 share?

17 benchmarks have published results for both models. Grok-2 (Dec 2024) has 34 scored results on Noometry and MiniMax M1 has 18.

Related comparisons

Go deeper