Model comparison

Llama 3-70B vs MiniMax M1

MiniMax M1 is the stronger model overall, scoring 40.3 to 28.8 on the Noometry Index.

Last verified . 17 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

MiniMax M1 MiniMax

40.3

Rank #150 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 3-70B scores higher in 0 categories and MiniMax M1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where MiniMax M1 leads 37.5 to 12.8.

Side by side

Llama 3-70B and MiniMax M1 specifications
Llama 3-70BMiniMax M1
ProviderMetaMiniMax
Noometry Index28.840.3
Released2024-04-182025-06-13
WeightsOpenOpen
Context window—1M
Max output—40K
Input $ / M tokens—$0.55
Output $ / M tokens—$2.20
Results tracked3118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax M1 leads

Llama 3-70B: 35.8 (#218), MiniMax M1: 39.9 (#153)

Coding benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Coding12061359
BigCodeBench Instruct43.6%—
BigCodeBench Complete54.5%—
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use Not comparable

Llama 3-70B: 21.1 (#139), MiniMax M1: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BMiniMax M1
Cybench5%—

Reasoning MiniMax M1 leads

Llama 3-70B: 18.0 (#288), MiniMax M1: 26.9 (#126)

Reasoning benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Hard Prompts11951339
Kagi LLM Benchmark35.1%—
DTBench54.2%—
Epoch Capabilities Index122.93—
ForecastBench57.1—
WinoGrande83.5%—

Math MiniMax M1 leads

Llama 3-70B: 12.8 (#305), MiniMax M1: 37.5 (#151)

Math benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Math12181361
OTIS Mock AIME 2024-20254.3%—
MATH Level 522.6%—

Knowledge MiniMax M1 leads

Llama 3-70B: 20.8 (#277), MiniMax M1: 36.4 (#170)

Knowledge benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Expert11491317
GPQA Diamond40.6%—
MMLU79.3%—

Multilingual MiniMax M1 leads

Llama 3-70B: 33.6 (#251), MiniMax M1: 45.8 (#163)

Multilingual benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Non-English11421319
LMArena Chinese11141360
LMArena French12321370
LMArena German11691350
LMArena Japanese10171217
LMArena Korean10171266
LMArena Russian11591329
LMArena Spanish12411353

Instruction Following MiniMax M1 leads

Llama 3-70B: 62.5 (#238), MiniMax M1: 69.3 (#174)

Instruction Following benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Instruction Following11941312

Long Context MiniMax M1 leads

Llama 3-70B: 35.6 (#240), MiniMax M1: 41.4 (#141)

Long Context benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Longer Query11741326
Fiction.LiveBench—69.4%

Writing & Preference MiniMax M1 leads

Llama 3-70B: 42.8 (#231), MiniMax M1: 53.1 (#161)

Writing & Preference benchmarks
BenchmarkLlama 3-70BMiniMax M1
LMArena Text12211343
LMArena Creative Writing12101298
LMArena Multi-Turn12231335

Frequently asked questions

Is Llama 3-70B better than MiniMax M1?

MiniMax M1 is the stronger model overall, scoring 40.3 to 28.8 on the Noometry Index.

Is Llama 3-70B or MiniMax M1 better for coding?

MiniMax M1 scores higher on coding benchmarks: 39.9 versus 35.8 in the Noometry coding category.

How many benchmarks do Llama 3-70B and MiniMax M1 share?

17 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and MiniMax M1 has 18.

Related comparisons

Go deeper