Model comparison

Llama 3.1-70B vs Mistral

Llama 3.1-70B and Mistral score almost the same on the Noometry Index (29.6 vs 29.9), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Llama 3.1-70B scores higher in 4 categories and Mistral in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.1-70B leads 65.3 to 52.6.
  • The biggest single-benchmark swing is MMLU-Pro: 65.3% for Llama 3.1-70B and 27.7% for Mistral.
  • Llama 3.1-70B has downloadable open weights; the other is API-only.

Side by side

Llama 3.1-70B and Mistral specifications
Llama 3.1-70BMistral
ProviderMetaMistral AI
Noometry Index29.629.9
Released2024-07-23—
WeightsOpenProprietary
Context window128K—
Max output4K—
Input $ / M tokens$0.40—
Output $ / M tokens$0.40—
Results tracked3522

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Llama 3.1-70B: 30.3 (#296), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkLlama 3.1-70BMistral
LMArena Coding12601162
WeirdML9%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete54.8%—

Agentic & Tool Use Not comparable

Llama 3.1-70B: 25.1 (#112), Mistral: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-70BMistral
TheAgentCompany6.9%—
BALROG27.9%—

Reasoning Too close to call

Llama 3.1-70B: 21.6 (#220), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkLlama 3.1-70BMistral
LMArena Hard Prompts12411149
DTBench60%—
LMCA14.8%—
Epoch Capabilities Index125.92—

Math Mistral leads

Llama 3.1-70B: 13.5 (#304), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkLlama 3.1-70BMistral
Omni-MATH21%7.2%
LMArena Math12521180
OTIS Mock AIME 2024-20253.6%—
MATH Level 536.7%—

Knowledge Llama 3.1-70B leads

Llama 3.1-70B: 24.2 (#269), Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkLlama 3.1-70BMistral
MMLU-Pro65.3%27.7%
GPQA (HELM)42.6%30.3%
LMArena Expert12091125
GPQA Diamond44.2%—
MMLU80.1%—

Multilingual Llama 3.1-70B leads

Llama 3.1-70B: 38.8 (#225), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkLlama 3.1-70BMistral
LMArena Non-English12191129
LMArena Chinese12151109
LMArena French12611180
LMArena German12221155
LMArena Japanese11321013
LMArena Korean11401032
LMArena Russian12341168
LMArena Spanish12531143

Instruction Following Llama 3.1-70B leads

Llama 3.1-70B: 65.3 (#223), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkLlama 3.1-70BMistral
IFEval82.1%56.8%
LMArena Instruction Following12311152

Long Context Llama 3.1-70B leads

Llama 3.1-70B: 37.6 (#214), Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkLlama 3.1-70BMistral
LMArena Longer Query12411153

Writing & Preference Mistral leads

Llama 3.1-70B: 35.4 (#267), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkLlama 3.1-70BMistral
LMArena Text12611165
LMArena Creative Writing12321158
WildBench75.8%66%
LMArena Multi-Turn12561147
EQ-Bench Creative Writing784—

Frequently asked questions

Is Llama 3.1-70B better than Mistral?

Llama 3.1-70B and Mistral score almost the same on the Noometry Index (29.6 vs 29.9), so choose on price, context window or the category you care about most.

Is Llama 3.1-70B or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 30.3 in the Noometry coding category.

How many benchmarks do Llama 3.1-70B and Mistral share?

22 benchmarks have published results for both models. Llama 3.1-70B has 35 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper