Model comparison

Llama 3.1-405B vs Mixtral 8x22B

Llama 3.1-405B is the stronger model overall, scoring 30.7 to 27.1 on the Noometry Index.

Last verified . 30 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Llama 3.1-405B scores higher in 6 categories and Mixtral 8x22B in 3 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3.1-405B leads 30.4 to 15.1.
  • The biggest single-benchmark swing is MMLU-Pro: 72.3% for Llama 3.1-405B and 46% for Mixtral 8x22B.

Side by side

Llama 3.1-405B and Mixtral 8x22B specifications
Llama 3.1-405BMixtral 8x22B
ProviderMetaMistral AI
Noometry Index30.727.1
Released2024-07-232024-04-17
WeightsOpenOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked4234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Llama 3.1-405B: 33.1 (#262), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
WeirdML21.4%3.2%
LMArena Coding12911166
BigCodeBench Instruct—40.6%
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Mixtral 8x22B leads

Llama 3.1-405B: 21.0 (#140), Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
Cybench7.5%7.5%
TheAgentCompany7.4%—

Reasoning Mixtral 8x22B leads

Llama 3.1-405B: 16.8 (#300), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
LMArena Hard Prompts12691150
DTBench61.4%55.1%
Epoch Capabilities Index128.75122.03
ForecastBench59.956.3
SimpleBench23%—
Kagi LLM Benchmark45%—
BIG-Bench Hard82.9%—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Mixtral 8x22B leads

Llama 3.1-405B: 18.4 (#290), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
Omni-MATH24.9%16.3%
LMArena Math12811184
MATH Level 549.8%24.2%
OTIS Mock AIME 2024-20259.7%—

Knowledge Llama 3.1-405B leads

Llama 3.1-405B: 30.4 (#227), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
GPQA Diamond50.9%34.1%
MMLU-Pro72.3%46%
GPQA (HELM)52.2%33.4%
LMArena Expert12431113
MMLU84.5%77.8%
Confabulations17.6%—
ARC (AI2) Challenge95.3%—
TriviaQA82.7%—

Multilingual Llama 3.1-405B leads

Llama 3.1-405B: 40.7 (#214), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
LMArena Non-English12481128
LMArena Chinese12421116
LMArena French12791166
LMArena German12521141
LMArena Japanese12081037
LMArena Korean11841057
LMArena Russian12651158
LMArena Spanish12601151

Instruction Following Llama 3.1-405B leads

Llama 3.1-405B: 65.9 (#214), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
IFEval81.1%72.4%
LMArena Instruction Following12591147

Long Context Llama 3.1-405B leads

Llama 3.1-405B: 38.4 (#197), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
LMArena Longer Query12661144

Writing & Preference Llama 3.1-405B leads

Llama 3.1-405B: 38.9 (#251), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BMixtral 8x22B
LMArena Text12841162
LMArena Creative Writing12621141
WildBench78.3%71.1%
LMArena Multi-Turn12971130
EQ-Bench Creative Writing870—

Frequently asked questions

Is Llama 3.1-405B better than Mixtral 8x22B?

Llama 3.1-405B is the stronger model overall, scoring 30.7 to 27.1 on the Noometry Index.

Is Llama 3.1-405B or Mixtral 8x22B better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 24.2 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Mixtral 8x22B share?

30 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper