Model comparison

Llama 3-70B vs MiniMax-M2.5

MiniMax-M2.5 is the stronger model overall, scoring 38.3 to 28.8 on the Noometry Index.

Last verified . 19 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

MiniMax-M2.5 MiniMax

38.3

Rank #188 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Llama 3-70B scores higher in 1 category and MiniMax-M2.5 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiniMax-M2.5 leads 39.2 to 20.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 35.1% for Llama 3-70B and 55.2% for MiniMax-M2.5.

Side by side

Llama 3-70B and MiniMax-M2.5 specifications
Llama 3-70BMiniMax-M2.5
ProviderMetaMiniMax
Noometry Index28.838.3
Released2024-04-182026-02-12
WeightsOpenOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked3133

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.5 leads

Llama 3-70B: 35.8 (#218), MiniMax-M2.5: 48.1 (#58)

Coding benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Coding12061381
SWE-bench Verified (bash only)—75.8%
LMArena WebDev—1387
SWE-bench Multilingual—68.3%
BigCodeBench Instruct43.6%—
BigCodeBench Complete54.5%—
ALE-Bench—618.17
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use MiniMax-M2.5 leads

Llama 3-70B: 21.1 (#139), MiniMax-M2.5: 30.4 (#77)

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
Terminal-Bench—42.7%
Cybench5%—
Vending-Bench 2—-23.16

Reasoning Too close to call

Llama 3-70B: 18.0 (#288), MiniMax-M2.5: 17.5 (#292)

Reasoning benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
Kagi LLM Benchmark35.1%55.2%
LMArena Hard Prompts11951372
Epoch Capabilities Index122.93146.68
ARC-AGI-2—4.9%
NYT Connections (extended)—16.8%
ARC-AGI-1—63.7%
DTBench54.2%—
ForecastBench57.1—
WinoGrande83.5%—

Math MiniMax-M2.5 leads

Llama 3-70B: 12.8 (#305), MiniMax-M2.5: 26.9 (#253)

Math benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Math12181378
OTIS Mock AIME 2024-20254.3%—
ProofBench—4%
MATH Level 522.6%—

Knowledge MiniMax-M2.5 leads

Llama 3-70B: 20.8 (#277), MiniMax-M2.5: 39.2 (#135)

Knowledge benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Expert11491379
GPQA Diamond40.6%—
Vectara Hallucination Rate—9.1%
MMLU79.3%—

Multilingual MiniMax-M2.5 leads

Llama 3-70B: 33.6 (#251), MiniMax-M2.5: 47.1 (#152)

Multilingual benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Non-English11421338
LMArena Chinese11141393
LMArena French12321362
LMArena German11691362
LMArena Japanese10171171
LMArena Korean10171232
LMArena Russian11591358
LMArena Spanish12411354

Instruction Following MiniMax-M2.5 leads

Llama 3-70B: 62.5 (#238), MiniMax-M2.5: 71.5 (#148)

Instruction Following benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Instruction Following11941353

Long Context MiniMax-M2.5 leads

Llama 3-70B: 35.6 (#240), MiniMax-M2.5: 37.5 (#216)

Long Context benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Longer Query11741366
CL-bench—11.4%
CL-bench Life—6.3%

Writing & Preference MiniMax-M2.5 leads

Llama 3-70B: 42.8 (#231), MiniMax-M2.5: 53.9 (#153)

Writing & Preference benchmarks
BenchmarkLlama 3-70BMiniMax-M2.5
LMArena Text12211359
LMArena Creative Writing12101331
LMArena Multi-Turn12231364
EQ-Bench Creative Writing—1361

Frequently asked questions

Is Llama 3-70B better than MiniMax-M2.5?

MiniMax-M2.5 is the stronger model overall, scoring 38.3 to 28.8 on the Noometry Index.

Is Llama 3-70B or MiniMax-M2.5 better for coding?

MiniMax-M2.5 scores higher on coding benchmarks: 48.1 versus 35.8 in the Noometry coding category.

How many benchmarks do Llama 3-70B and MiniMax-M2.5 share?

19 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and MiniMax-M2.5 has 33.

Related comparisons

Go deeper