Model comparison

Llama 3.1-70B vs Mixtral 8x22B

Llama 3.1-70B is the stronger model overall, scoring 29.6 to 27.1 on the Noometry Index.

Last verified . 30 shared benchmarks.

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Llama 3.1-70B scores higher in 7 categories and Mixtral 8x22B in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mixtral 8x22B leads 22.9 to 13.5.
  • The biggest single-benchmark swing is MMLU-Pro: 65.3% for Llama 3.1-70B and 46% for Mixtral 8x22B.
  • Llama 3.1-70B is cheaper at $0.40 / $0.40 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Llama 3.1-70B accepts more context: 128K tokens versus 64K.

Side by side

Llama 3.1-70B and Mixtral 8x22B specifications
Llama 3.1-70BMixtral 8x22B
ProviderMetaMistral AI
Noometry Index29.627.1
Released2024-07-232024-04-17
WeightsOpenOpen
Context window128K64K
Max output4K64K
Input $ / M tokens$0.40$2
Output $ / M tokens$0.40$6
Results tracked3534

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-70B leads

Llama 3.1-70B: 30.3 (#296), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
WeirdML9%3.2%
BigCodeBench Instruct46.1%40.6%
LMArena Coding12601166
BigCodeBench Complete54.8%50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Llama 3.1-70B leads

Llama 3.1-70B: 25.1 (#112), Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
TheAgentCompany6.9%—
Cybench—7.5%
BALROG27.9%—

Reasoning Llama 3.1-70B leads

Llama 3.1-70B: 21.6 (#220), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
LMArena Hard Prompts12411150
DTBench60%55.1%
Epoch Capabilities Index125.92122.03
LMCA14.8%—
ForecastBench—56.3

Math Mixtral 8x22B leads

Llama 3.1-70B: 13.5 (#304), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
Omni-MATH21%16.3%
LMArena Math12521184
MATH Level 536.7%24.2%
OTIS Mock AIME 2024-20253.6%—

Knowledge Llama 3.1-70B leads

Llama 3.1-70B: 24.2 (#269), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
GPQA Diamond44.2%34.1%
MMLU-Pro65.3%46%
GPQA (HELM)42.6%33.4%
LMArena Expert12091113
MMLU80.1%77.8%

Multilingual Llama 3.1-70B leads

Llama 3.1-70B: 38.8 (#225), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
LMArena Non-English12191128
LMArena Chinese12151116
LMArena French12611166
LMArena German12221141
LMArena Japanese11321037
LMArena Korean11401057
LMArena Russian12341158
LMArena Spanish12531151

Instruction Following Llama 3.1-70B leads

Llama 3.1-70B: 65.3 (#223), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
IFEval82.1%72.4%
LMArena Instruction Following12311147

Long Context Llama 3.1-70B leads

Llama 3.1-70B: 37.6 (#214), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
LMArena Longer Query12411144

Writing & Preference Mixtral 8x22B leads

Llama 3.1-70B: 35.4 (#267), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkLlama 3.1-70BMixtral 8x22B
LMArena Text12611162
LMArena Creative Writing12321141
WildBench75.8%71.1%
LMArena Multi-Turn12561130
EQ-Bench Creative Writing784—

Frequently asked questions

Is Llama 3.1-70B better than Mixtral 8x22B?

Llama 3.1-70B is the stronger model overall, scoring 29.6 to 27.1 on the Noometry Index.

Which is cheaper, Llama 3.1-70B or Mixtral 8x22B?

Llama 3.1-70B is cheaper. It lists at $0.40 per million input tokens and $0.40 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Llama 3.1-70B or Mixtral 8x22B better for coding?

Llama 3.1-70B scores higher on coding benchmarks: 30.3 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Llama 3.1-70B does, with 128K tokens against 64K.

How many benchmarks do Llama 3.1-70B and Mixtral 8x22B share?

30 benchmarks have published results for both models. Llama 3.1-70B has 35 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper