Model comparison

Llama 3.1-405B vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 30.7 on the Noometry Index.

Last verified . 32 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Llama 3.1-405B scores higher in 5 categories and Mistral Large in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Mistral Large leads 28.6 to 21.0.
  • The biggest single-benchmark swing is MMLU-Pro: 72.3% for Llama 3.1-405B and 59.9% for Mistral Large.

Side by side

Llama 3.1-405B and Mistral Large specifications
Llama 3.1-405BMistral Large
ProviderMetaMistral AI
Noometry Index30.731.9
Released2024-07-232024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked4251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Llama 3.1-405B: 33.1 (#262), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkLlama 3.1-405BMistral Large
LMArena Coding12911277
SciCode—36.2%
WeirdML21.4%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Llama 3.1-405B: 21.0 (#140), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BMistral Large
Berkeley Function Calling Leaderboard—38.4%
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Llama 3.1-405B leads

Llama 3.1-405B: 16.8 (#300), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkLlama 3.1-405BMistral Large
SimpleBench23%22.5%
LMArena Hard Prompts12691257
DTBench61.4%65.1%
Epoch Capabilities Index128.75128.52
ForecastBench59.957.1
Kagi LLM Benchmark45%—
CritPt—0%
LiveBench Reasoning—43.5%
LiveBench Data Analysis—50.1%
LMCA—16.7%
BIG-Bench Hard82.9%—
HellaSwag89.2%—
LiveBench—48.4%
PIQA85.9%—
WinoGrande89.2%—

Math Too close to call

Llama 3.1-405B: 18.4 (#290), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkLlama 3.1-405BMistral Large
OTIS Mock AIME 2024-20259.7%8.5%
Omni-MATH24.9%28.1%
LMArena Math12811262
MATH Level 549.8%50.3%
LiveBench Math—42.5%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Too close to call

Llama 3.1-405B: 30.4 (#227), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkLlama 3.1-405BMistral Large
GPQA Diamond50.9%51.3%
MMLU-Pro72.3%59.9%
Confabulations17.6%21.4%
GPQA (HELM)52.2%43.5%
LMArena Expert12431232
MMLU84.5%80%
Vectara Hallucination Rate—4.5%
ARC (AI2) Challenge95.3%—
TriviaQA82.7%—

Multilingual Too close to call

Llama 3.1-405B: 40.7 (#214), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkLlama 3.1-405BMistral Large
LMArena Non-English12481237
LMArena Chinese12421240
LMArena French12791325
LMArena German12521254
LMArena Japanese12081188
LMArena Korean11841202
LMArena Russian12651257
LMArena Spanish12601268

Instruction Following Mistral Large leads

Llama 3.1-405B: 65.9 (#214), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BMistral Large
IFEval81.1%87.7%
LMArena Instruction Following12591249
LiveBench Instruction Following—67.9%

Long Context Too close to call

Llama 3.1-405B: 38.4 (#197), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkLlama 3.1-405BMistral Large
LMArena Longer Query12661261

Writing & Preference Mistral Large leads

Llama 3.1-405B: 38.9 (#251), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BMistral Large
LMArena Text12841266
LMArena Creative Writing12621243
EQ-Bench Creative Writing870985
WildBench78.3%80.1%
LMArena Multi-Turn12971260
Short-Story Creative Writing—69%
LiveBench Language—39.4%

Frequently asked questions

Is Llama 3.1-405B better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 30.7 on the Noometry Index.

Is Llama 3.1-405B or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 33.1 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Mistral Large share?

32 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper