Model comparison

Llama2 70b Steerlm Chat vs Mixtral 8x22B

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 27.1 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 3 categories and Mixtral 8x22B in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 22.9.

Side by side

Llama2 70b Steerlm Chat and Mixtral 8x22B specifications
Llama2 70b Steerlm ChatMixtral 8x22B
ProviderNVIDIAMistral AI
Noometry Index31.827.1
Released—2024-04-17
WeightsOpenOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 29.9 (#300), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Coding10251166
WeirdML—3.2%
BigCodeBench Instruct—40.6%
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Not comparable

Llama2 70b Steerlm Chat: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
Cybench—7.5%

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Hard Prompts10471150
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.3 (#226), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Math10721184
Omni-MATH—16.3%
MATH Level 5—24.2%

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Mixtral 8x22B leads

Llama2 70b Steerlm Chat: 28.8 (#270), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Non-English10631128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Mixtral 8x22B leads

Llama2 70b Steerlm Chat: 54.2 (#279), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Instruction Following10601147
IFEval—72.4%

Long Context Mixtral 8x22B leads

Llama2 70b Steerlm Chat: 30.4 (#288), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Longer Query9981144

Writing & Preference Mixtral 8x22B leads

Llama2 70b Steerlm Chat: 31.6 (#283), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatMixtral 8x22B
LMArena Text10981162
LMArena Creative Writing10911141
LMArena Multi-Turn10581130
WildBench—71.1%

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Mixtral 8x22B?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 27.1 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Mixtral 8x22B better for coding?

Llama2 70b Steerlm Chat scores higher on coding benchmarks: 29.9 versus 24.2 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Mixtral 8x22B share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper