Model comparison

Llama 3.2 90B vs MiniMax-M2

MiniMax-M2 is the stronger model overall, scoring 37.4 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Summary

  • The widest gap is in math, where MiniMax-M2 leads 37.3 to 11.1.

Side by side

Llama 3.2 90B and MiniMax-M2 specifications
Llama 3.2 90BMiniMax-M2
ProviderMetaMiniMax
Noometry Index27.537.4
Released2024-09-242025-10-27
WeightsOpenOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked921

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, MiniMax-M2: 39.3 (#159)

Coding benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
SWE-bench Verified (bash only)—61%
LMArena WebDev—1297
LMArena Coding—1370

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), MiniMax-M2: 25.1 (#109)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
Terminal-Bench—30%
BALROG27.3%—
Vending-Bench 2—160.6

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), MiniMax-M2: 19.4 (#258)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
Kagi LLM Benchmark—57.8%
NYT Connections (extended)—14.8%
EnigmaEval0.4%—
LMArena Hard Prompts—1357
Epoch Capabilities Index125.5—

Math MiniMax-M2 leads

Llama 3.2 90B: 11.1 (#308), MiniMax-M2: 37.3 (#160)

Math benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
OTIS Mock AIME 2024-20252.6%—
LMArena Math—1352
MATH Level 539.4%—

Knowledge MiniMax-M2 leads

Llama 3.2 90B: 21.7 (#274), MiniMax-M2: 37.0 (#163)

Knowledge benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
GPQA Diamond41%—
LMArena Expert—1337
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), MiniMax-M2: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, MiniMax-M2: 45.3 (#171)

Multilingual benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
LMArena Non-English—1313
LMArena Chinese—1366
LMArena French—1335
LMArena German—1355
LMArena Russian—1331
LMArena Spanish—1326

Instruction Following Not comparable

Llama 3.2 90B: —, MiniMax-M2: 70.2 (#166)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
LMArena Instruction Following—1328

Long Context Not comparable

Llama 3.2 90B: —, MiniMax-M2: 40.5 (#153)

Long Context benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
LMArena Longer Query—1331

Writing & Preference Not comparable

Llama 3.2 90B: —, MiniMax-M2: 53.0 (#162)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BMiniMax-M2
LMArena Text—1340
LMArena Creative Writing—1286
LMArena Multi-Turn—1361

Frequently asked questions

Is Llama 3.2 90B better than MiniMax-M2?

MiniMax-M2 is the stronger model overall, scoring 37.4 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and MiniMax-M2 share?

0 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and MiniMax-M2 has 21.

Related comparisons

Go deeper