Model comparison

Llama 3.1-405B vs Mistral Small 3.2

Llama 3.1-405B and Mistral Small 3.2 score almost the same on the Noometry Index (30.7 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Mistral Small 3.2 Mistral AI

31.2

Rank #280 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Llama 3.1-405B scores higher in 1 category and Mistral Small 3.2 in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Small 3.2 leads 26.3 to 18.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 9.7% for Llama 3.1-405B and 30.3% for Mistral Small 3.2.

Side by side

Llama 3.1-405B and Mistral Small 3.2 specifications
Llama 3.1-405BMistral Small 3.2
ProviderMetaMistral AI
Noometry Index30.731.2
Released2024-07-232025-06-20
WeightsOpenOpen
Context window—256K
Max output—16K
Input $ / M tokens—$0.0938
Output $ / M tokens—$0.25
Results tracked426

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.1-405B: 33.1 (#262), Mistral Small 3.2: —

Coding benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
WeirdML21.4%—
LMArena Coding1291—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), Mistral Small 3.2: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Mistral Small 3.2 leads

Llama 3.1-405B: 16.8 (#300), Mistral Small 3.2: 18.1 (#287)

Reasoning benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
Kagi LLM Benchmark45%40.4%
Epoch Capabilities Index128.75131.74
SimpleBench23%—
Chess Puzzles—1%
LMArena Hard Prompts1269—
DTBench61.4%—
BIG-Bench Hard82.9%—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Mistral Small 3.2 leads

Llama 3.1-405B: 18.4 (#290), Mistral Small 3.2: 26.3 (#260)

Math benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
OTIS Mock AIME 2024-20259.7%30.3%
Omni-MATH24.9%—
LMArena Math1281—
MATH Level 549.8%—

Knowledge Llama 3.1-405B leads

Llama 3.1-405B: 30.4 (#227), Mistral Small 3.2: 26.7 (#256)

Knowledge benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
GPQA Diamond50.9%49.1%
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
LMArena Expert1243—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Not comparable

Llama 3.1-405B: 40.7 (#214), Mistral Small 3.2: —

Multilingual benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
LMArena Non-English1248—
LMArena Chinese1242—
LMArena French1279—
LMArena German1252—
LMArena Japanese1208—
LMArena Korean1184—
LMArena Russian1265—
LMArena Spanish1260—

Instruction Following Not comparable

Llama 3.1-405B: 65.9 (#214), Mistral Small 3.2: —

Instruction Following benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
IFEval81.1%—
LMArena Instruction Following1259—

Long Context Not comparable

Llama 3.1-405B: 38.4 (#197), Mistral Small 3.2: —

Long Context benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
LMArena Longer Query1266—

Writing & Preference Mistral Small 3.2 leads

Llama 3.1-405B: 38.9 (#251), Mistral Small 3.2: 45.0 (#224)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BMistral Small 3.2
EQ-Bench Creative Writing8701255
LMArena Text1284—
LMArena Creative Writing1262—
WildBench78.3%—
LMArena Multi-Turn1297—

Frequently asked questions

Is Llama 3.1-405B better than Mistral Small 3.2?

Llama 3.1-405B and Mistral Small 3.2 score almost the same on the Noometry Index (30.7 vs 31.2), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3.1-405B and Mistral Small 3.2 share?

5 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Mistral Small 3.2 has 6.

Related comparisons

Go deeper