Model comparison

Llama 3-8B vs Mixtral 8x7B

Mixtral 8x7B is the stronger model overall, scoring 27.1 to 25.5 on the Noometry Index.

Last verified . 30 shared benchmarks.

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Llama 3-8B scores higher in 4 categories and Mixtral 8x7B in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mixtral 8x7B leads 18.8 to 8.8.
  • The biggest single-benchmark swing is DTBench: 43.9% for Llama 3-8B and 49.6% for Mixtral 8x7B.

Side by side

Llama 3-8B and Mixtral 8x7B specifications
Llama 3-8BMixtral 8x7B
ProviderMetaMistral AI
Noometry Index25.527.1
Released2024-04-182023-12-11
WeightsOpenOpen
Context window—32K
Max output—32K
Input $ / M tokens—$0.70
Output $ / M tokens—$0.70
Results tracked3438

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mixtral 8x7B leads

Llama 3-8B: 31.0 (#289), Mixtral 8x7B: 32.8 (#269)

Coding benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Coding11521126
HumanEval+56.7%39.6%
MBPP+54.8%49.7%
BigCodeBench Instruct31.9%—
BigCodeBench Complete36.9%—

Reasoning Mixtral 8x7B leads

Llama 3-8B: 14.3 (#326), Mixtral 8x7B: 18.2 (#285)

Reasoning benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Hard Prompts11331115
DTBench43.9%49.6%
Adversarial NLI57.3%55.2%
Epoch Capabilities Index116.45118.47
ForecastBench58.656.3
WinoGrande75.7%77.2%
Chess Puzzles0%—
HellaSwag—86.7%
PIQA—83.6%

Math Mixtral 8x7B leads

Llama 3-8B: 8.8 (#323), Mixtral 8x7B: 18.8 (#289)

Math benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Math11511147
MATH Level 56.1%10%
OTIS Mock AIME 2024-20251.9%—
Omni-MATH—10.5%
GSM8K—74.4%

Knowledge Mixtral 8x7B leads

Llama 3-8B: 7.8 (#308), Mixtral 8x7B: 11.0 (#301)

Knowledge benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
GPQA Diamond26.1%30.6%
LMArena Expert11131088
ARC (AI2) Challenge82.8%87.3%
MMLU68.8%70.6%
OpenBookQA82.6%85.8%
TriviaQA67.7%82.2%
MMLU-Pro—33.5%
GPQA (HELM)—29.6%

Multilingual Llama 3-8B leads

Llama 3-8B: 30.8 (#261), Mixtral 8x7B: 29.6 (#266)

Multilingual benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Non-English10981077
LMArena Chinese10761055
LMArena French11591166
LMArena German11041114
LMArena Japanese967931
LMArena Korean1004968
LMArena Russian11091090
LMArena Spanish11731111

Instruction Following Llama 3-8B leads

Llama 3-8B: 58.4 (#260), Mixtral 8x7B: 51.0 (#297)

Instruction Following benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Instruction Following11271109
IFEval—57.5%

Long Context Too close to call

Llama 3-8B: 34.2 (#251), Mixtral 8x7B: 33.4 (#260)

Long Context benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Longer Query11281103

Writing & Preference Llama 3-8B leads

Llama 3-8B: 37.5 (#256), Mixtral 8x7B: 34.2 (#270)

Writing & Preference benchmarks
BenchmarkLlama 3-8BMixtral 8x7B
LMArena Text11661132
LMArena Creative Writing11501109
LMArena Multi-Turn11521115
WildBench—67.3%

Frequently asked questions

Is Llama 3-8B better than Mixtral 8x7B?

Mixtral 8x7B is the stronger model overall, scoring 27.1 to 25.5 on the Noometry Index.

Is Llama 3-8B or Mixtral 8x7B better for coding?

Mixtral 8x7B scores higher on coding benchmarks: 32.8 versus 31.0 in the Noometry coding category.

How many benchmarks do Llama 3-8B and Mixtral 8x7B share?

30 benchmarks have published results for both models. Llama 3-8B has 34 scored results on Noometry and Mixtral 8x7B has 38.

Related comparisons

Go deeper