Model comparison

Llama 3.1-405B vs Mistral

Llama 3.1-405B and Mistral score almost the same on the Noometry Index (30.7 vs 29.9), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Llama 3.1-405B scores higher in 5 categories and Mistral in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3.1-405B leads 30.4 to 16.6.
  • The biggest single-benchmark swing is MMLU-Pro: 72.3% for Llama 3.1-405B and 27.7% for Mistral.
  • Llama 3.1-405B has downloadable open weights; the other is API-only.

Side by side

Llama 3.1-405B and Mistral specifications
Llama 3.1-405BMistral
ProviderMetaMistral AI
Noometry Index30.729.9
Released2024-07-23—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4222

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.1-405B: 33.1 (#262), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkLlama 3.1-405BMistral
LMArena Coding12911162
WeirdML21.4%—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), Mistral: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BMistral
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Mistral leads

Llama 3.1-405B: 16.8 (#300), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkLlama 3.1-405BMistral
LMArena Hard Prompts12691149
SimpleBench23%—
Kagi LLM Benchmark45%—
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Mistral leads

Llama 3.1-405B: 18.4 (#290), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkLlama 3.1-405BMistral
Omni-MATH24.9%7.2%
LMArena Math12811180
OTIS Mock AIME 2024-20259.7%—
MATH Level 549.8%—

Knowledge Llama 3.1-405B leads

Llama 3.1-405B: 30.4 (#227), Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkLlama 3.1-405BMistral
MMLU-Pro72.3%27.7%
GPQA (HELM)52.2%30.3%
LMArena Expert12431125
GPQA Diamond50.9%—
Confabulations17.6%—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Llama 3.1-405B leads

Llama 3.1-405B: 40.7 (#214), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkLlama 3.1-405BMistral
LMArena Non-English12481129
LMArena Chinese12421109
LMArena French12791180
LMArena German12521155
LMArena Japanese12081013
LMArena Korean11841032
LMArena Russian12651168
LMArena Spanish12601143

Instruction Following Llama 3.1-405B leads

Llama 3.1-405B: 65.9 (#214), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BMistral
IFEval81.1%56.8%
LMArena Instruction Following12591152

Long Context Llama 3.1-405B leads

Llama 3.1-405B: 38.4 (#197), Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkLlama 3.1-405BMistral
LMArena Longer Query12661153

Writing & Preference Llama 3.1-405B leads

Llama 3.1-405B: 38.9 (#251), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BMistral
LMArena Text12841165
LMArena Creative Writing12621158
WildBench78.3%66%
LMArena Multi-Turn12971147
EQ-Bench Creative Writing870—

Frequently asked questions

Is Llama 3.1-405B better than Mistral?

Llama 3.1-405B and Mistral score almost the same on the Noometry Index (30.7 vs 29.9), so choose on price, context window or the category you care about most.

Is Llama 3.1-405B or Mistral better for coding?

They score almost the same on coding (33.1 vs 33.8); test both on your own repository before choosing.

How many benchmarks do Llama 3.1-405B and Mistral share?

22 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper