Model comparison

Llama2 70b Steerlm Chat vs Mistral

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 29.9 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 2 categories and Mistral in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 22.3.
  • Llama2 70b Steerlm Chat has downloadable open weights; the other is API-only.

Side by side

Llama2 70b Steerlm Chat and Mistral specifications
Llama2 70b Steerlm ChatMistral
ProviderNVIDIAMistral AI
Noometry Index31.829.9
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked922

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Llama2 70b Steerlm Chat: 29.9 (#300), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Coding10251162

Reasoning Mistral leads

Llama2 70b Steerlm Chat: 20.0 (#246), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Hard Prompts10471149

Math Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.3 (#226), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Math10721180
Omni-MATH—7.2%

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
LMArena Expert—1125

Multilingual Mistral leads

Llama2 70b Steerlm Chat: 28.8 (#270), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Non-English10631129
LMArena Chinese—1109
LMArena French—1180
LMArena German—1155
LMArena Japanese—1013
LMArena Korean—1032
LMArena Russian—1168
LMArena Spanish—1143

Instruction Following Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 54.2 (#279), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Instruction Following10601152
IFEval—56.8%

Long Context Mistral leads

Llama2 70b Steerlm Chat: 30.4 (#288), Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Longer Query9981153

Writing & Preference Mistral leads

Llama2 70b Steerlm Chat: 31.6 (#283), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral
LMArena Text10981165
LMArena Creative Writing10911158
LMArena Multi-Turn10581147
WildBench—66%

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Mistral?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 29.9 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Mistral share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper