Model comparison

Llama-3.3-70B-Instruct vs MiniMax-M2.1

MiniMax-M2.1 is the stronger model overall, scoring 38.9 to 30.6 on the Noometry Index. Llama-3.3-70B-Instruct costs 3.4× less per token, which makes it the better buy when MiniMax-M2.1's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

MiniMax-M2.1 MiniMax

38.9

Rank #178 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Llama-3.3-70B-Instruct scores higher in 0 categories and MiniMax-M2.1 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where MiniMax-M2.1 leads 38.3 to 15.3.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 4.1% for Llama-3.3-70B-Instruct and 11.8% for MiniMax-M2.1.
  • Llama-3.3-70B-Instruct is cheaper at $0.10 / $0.32 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.1.
  • MiniMax-M2.1 accepts more context: 205K tokens versus 128K.

Side by side

Llama-3.3-70B-Instruct and MiniMax-M2.1 specifications
Llama-3.3-70B-InstructMiniMax-M2.1
ProviderMetaMiniMax
Noometry Index30.638.9
Released2024-12-062025-12-23
WeightsOpenOpen
Context window128K205K
Max output4K131K
Input $ / M tokens$0.10$0.30
Output $ / M tokens$0.32$1.20
Results tracked4322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 31.0 (#290), MiniMax-M2.1: 40.4 (#143)

Coding benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Coding12681421
LMArena WebDev—1384
SciCode26%—
WeirdML14.4%—
BigCodeBench Instruct46.9%—
LiveBench Coding36.6%—
BigCodeBench Complete57.5%—
ALE-Bench—623.83

Agentic & Tool Use MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 25.8 (#105), MiniMax-M2.1: 27.9 (#98)

Agentic & Tool Use benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
Terminal-Bench—36.6%
Berkeley Function Calling Leaderboard31.9%—
BALROG23%—

Reasoning MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 14.1 (#327), MiniMax-M2.1: 16.6 (#302)

Reasoning benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Hard Prompts12571411
SimpleBench19.9%—
NYT Connections (extended)—11.2%
CritPt0%—
LiveBench Reasoning50.8%—
DTBench59.5%—
LiveBench Data Analysis49.5%—
LMCA17.5%—
Epoch Capabilities Index127.33—
ForecastBench58.6—
LiveBench50.2%—

Math MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 15.3 (#298), MiniMax-M2.1: 38.3 (#138)

Math benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Math12671397
OTIS Mock AIME 2024-20255.1%—
LiveBench Math42.2%—
MATH Level 541.6%—

Knowledge MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 30.6 (#226), MiniMax-M2.1: 38.3 (#147)

Knowledge benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
Vectara Hallucination Rate4.1%11.8%
LMArena Expert12251431
GPQA Diamond47.4%—
Confabulations22.8%—
MMLU86.3%—

Multilingual MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 39.9 (#220), MiniMax-M2.1: 50.0 (#128)

Multilingual benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Non-English12361378
LMArena Chinese12171430
LMArena French12811404
LMArena German12511381
LMArena Japanese11501287
LMArena Korean11431298
LMArena Russian12521387
LMArena Spanish12701397

Instruction Following MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 71.1 (#157), MiniMax-M2.1: 73.8 (#112)

Instruction Following benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Instruction Following12421400
LiveBench Instruction Following82.7%—

Long Context MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 26.4 (#295), MiniMax-M2.1: 43.2 (#101)

Long Context benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Longer Query12561416
Fiction.LiveBench33.3%—

Writing & Preference MiniMax-M2.1 leads

Llama-3.3-70B-Instruct: 47.6 (#207), MiniMax-M2.1: 58.3 (#120)

Writing & Preference benchmarks
BenchmarkLlama-3.3-70B-InstructMiniMax-M2.1
LMArena Text12741392
LMArena Creative Writing12501361
LMArena Multi-Turn12801396
LiveBench Language39.2%—

Frequently asked questions

Is Llama-3.3-70B-Instruct better than MiniMax-M2.1?

MiniMax-M2.1 is the stronger model overall, scoring 38.9 to 30.6 on the Noometry Index. Llama-3.3-70B-Instruct costs 3.4× less per token, which makes it the better buy when MiniMax-M2.1's lead doesn't matter for your workload.

Which is cheaper, Llama-3.3-70B-Instruct or MiniMax-M2.1?

Llama-3.3-70B-Instruct is cheaper. It lists at $0.10 per million input tokens and $0.32 per million output tokens; MiniMax-M2.1 lists at $0.30 and $1.20.

Is Llama-3.3-70B-Instruct or MiniMax-M2.1 better for coding?

MiniMax-M2.1 scores higher on coding benchmarks: 40.4 versus 31.0 in the Noometry coding category.

Which has the bigger context window?

MiniMax-M2.1 does, with 205K tokens against 128K.

How many benchmarks do Llama-3.3-70B-Instruct and MiniMax-M2.1 share?

18 benchmarks have published results for both models. Llama-3.3-70B-Instruct has 43 scored results on Noometry and MiniMax-M2.1 has 22.

Related comparisons

Go deeper