Model comparison
Llama 3.1-405B vs Mistral Small 3.2
Llama 3.1-405B and Mistral Small 3.2 score almost the same on the Noometry Index (30.7 vs 31.2), so choose on price, context window or the category you care about most.
Last verified . 5 shared benchmarks.
Summary
- They share 5 benchmarks with published results for both. Llama 3.1-405B scores higher in 1 category and Mistral Small 3.2 in 3 categories; 4 gaps are clear of the uncertainty.
- The widest gap is in math, where Mistral Small 3.2 leads 26.3 to 18.4.
- The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 9.7% for Llama 3.1-405B and 30.3% for Mistral Small 3.2.
Side by side
| Llama 3.1-405B | Mistral Small 3.2 | |
|---|---|---|
| Provider | Meta | Mistral AI |
| Noometry Index | 30.7 | 31.2 |
| Released | 2024-07-23 | 2025-06-20 |
| Weights | Open | Open |
| Context window | — | 256K |
| Max output | — | 16K |
| Input $ / M tokens | — | $0.0938 |
| Output $ / M tokens | — | $0.25 |
| Results tracked | 42 | 6 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Llama 3.1-405B: 33.1 (#262), Mistral Small 3.2: —
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| WeirdML | 21.4% | — |
| LMArena Coding | 1291 | — |
Agentic & Tool Use Not comparable
Llama 3.1-405B: 21.0 (#140), Mistral Small 3.2: —
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| TheAgentCompany | 7.4% | — |
| Cybench | 7.5% | — |
Reasoning Mistral Small 3.2 leads
Llama 3.1-405B: 16.8 (#300), Mistral Small 3.2: 18.1 (#287)
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| Kagi LLM Benchmark | 45% | 40.4% |
| Epoch Capabilities Index | 128.75 | 131.74 |
| SimpleBench | 23% | — |
| Chess Puzzles | — | 1% |
| LMArena Hard Prompts | 1269 | — |
| DTBench | 61.4% | — |
| BIG-Bench Hard | 82.9% | — |
| ForecastBench | 59.9 | — |
| HellaSwag | 89.2% | — |
| PIQA | 85.9% | — |
| WinoGrande | 89.2% | — |
Math Mistral Small 3.2 leads
Llama 3.1-405B: 18.4 (#290), Mistral Small 3.2: 26.3 (#260)
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 9.7% | 30.3% |
| Omni-MATH | 24.9% | — |
| LMArena Math | 1281 | — |
| MATH Level 5 | 49.8% | — |
Knowledge Llama 3.1-405B leads
Llama 3.1-405B: 30.4 (#227), Mistral Small 3.2: 26.7 (#256)
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| GPQA Diamond | 50.9% | 49.1% |
| MMLU-Pro | 72.3% | — |
| Confabulations | 17.6% | — |
| GPQA (HELM) | 52.2% | — |
| LMArena Expert | 1243 | — |
| ARC (AI2) Challenge | 95.3% | — |
| MMLU | 84.5% | — |
| TriviaQA | 82.7% | — |
Multilingual Not comparable
Llama 3.1-405B: 40.7 (#214), Mistral Small 3.2: —
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| LMArena Non-English | 1248 | — |
| LMArena Chinese | 1242 | — |
| LMArena French | 1279 | — |
| LMArena German | 1252 | — |
| LMArena Japanese | 1208 | — |
| LMArena Korean | 1184 | — |
| LMArena Russian | 1265 | — |
| LMArena Spanish | 1260 | — |
Instruction Following Not comparable
Llama 3.1-405B: 65.9 (#214), Mistral Small 3.2: —
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| IFEval | 81.1% | — |
| LMArena Instruction Following | 1259 | — |
Long Context Not comparable
Llama 3.1-405B: 38.4 (#197), Mistral Small 3.2: —
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| LMArena Longer Query | 1266 | — |
Writing & Preference Mistral Small 3.2 leads
Llama 3.1-405B: 38.9 (#251), Mistral Small 3.2: 45.0 (#224)
| Benchmark | Llama 3.1-405B | Mistral Small 3.2 |
|---|---|---|
| EQ-Bench Creative Writing | 870 | 1255 |
| LMArena Text | 1284 | — |
| LMArena Creative Writing | 1262 | — |
| WildBench | 78.3% | — |
| LMArena Multi-Turn | 1297 | — |
Frequently asked questions
Is Llama 3.1-405B better than Mistral Small 3.2?
Llama 3.1-405B and Mistral Small 3.2 score almost the same on the Noometry Index (30.7 vs 31.2), so choose on price, context window or the category you care about most.
How many benchmarks do Llama 3.1-405B and Mistral Small 3.2 share?
5 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Mistral Small 3.2 has 6.