Model comparison

Llama 3-70B vs Mistral Medium 3.5

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 28.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Llama 3-70B scores higher in 1 category and Mistral Medium 3.5 in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium 3.5 leads 39.1 to 12.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 35.1% for Llama 3-70B and 41.4% for Mistral Medium 3.5.

Side by side

Llama 3-70B and Mistral Medium 3.5 specifications
Llama 3-70BMistral Medium 3.5
ProviderMetaMistral AI
Noometry Index28.840.2
Released2024-04-18—
WeightsOpenOpen
Context window—262K
Max output—210K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked3122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3-70B: 35.8 (#218), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Coding12061461
LMArena WebDev—1264
BigCodeBench Instruct43.6%—
BigCodeBench Complete54.5%—
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use Not comparable

Llama 3-70B: 21.1 (#139), Mistral Medium 3.5: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
Cybench5%—

Reasoning Too close to call

Llama 3-70B: 18.0 (#288), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
Kagi LLM Benchmark35.1%41.4%
LMArena Hard Prompts11951436
Epoch Capabilities Index122.93141.35
NYT Connections (extended)—12.9%
DTBench54.2%—
ForecastBench57.1—
WinoGrande83.5%—

Math Mistral Medium 3.5 leads

Llama 3-70B: 12.8 (#305), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Math12181431
OTIS Mock AIME 2024-20254.3%—
MATH Level 522.6%—

Knowledge Mistral Medium 3.5 leads

Llama 3-70B: 20.8 (#277), Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Expert11491432
GPQA Diamond40.6%—
MMLU79.3%—

Multimodal Not comparable

Llama 3-70B: —, Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Vision—1223

Multilingual Mistral Medium 3.5 leads

Llama 3-70B: 33.6 (#251), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Non-English11421404
LMArena Chinese11141442
LMArena French12321448
LMArena German11691451
LMArena Korean10171385
LMArena Russian11591395
LMArena Spanish12411409
LMArena Japanese1017—

Instruction Following Mistral Medium 3.5 leads

Llama 3-70B: 62.5 (#238), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Instruction Following11941415

Long Context Mistral Medium 3.5 leads

Llama 3-70B: 35.6 (#240), Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Longer Query11741415

Writing & Preference Mistral Medium 3.5 leads

Llama 3-70B: 42.8 (#231), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkLlama 3-70BMistral Medium 3.5
LMArena Text12211421
LMArena Creative Writing12101374
LMArena Multi-Turn12231423
EQ-Bench 4—993

Frequently asked questions

Is Llama 3-70B better than Mistral Medium 3.5?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 28.8 on the Noometry Index.

Is Llama 3-70B or Mistral Medium 3.5 better for coding?

They score almost the same on coding (35.8 vs 36.0); test both on your own repository before choosing.

How many benchmarks do Llama 3-70B and Mistral Medium 3.5 share?

18 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper