Model comparison

Llama2 70b Steerlm Chat vs Mistral Nemo

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 26.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 25.5.

Side by side

Llama2 70b Steerlm Chat and Mistral Nemo specifications
Llama2 70b Steerlm ChatMistral Nemo
ProviderNVIDIAMistral AI
Noometry Index31.826.4
Released—2024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked910

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama2 70b Steerlm Chat: 29.9 (#300), Mistral Nemo: —

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Coding1025—

Agentic & Tool Use Not comparable

Llama2 70b Steerlm Chat: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Hard Prompts1047—
DTBench—48.6%
Epoch Capabilities Index—118.68
PIQA—83.5%

Math Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.3 (#226), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Math1072—
MATH Level 5—10.8%
GSM8K—84.2%

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
GPQA Diamond—29.9%
BoolQ—82.5%

Multilingual Not comparable

Llama2 70b Steerlm Chat: 28.8 (#270), Mistral Nemo: —

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Non-English1063—

Instruction Following Not comparable

Llama2 70b Steerlm Chat: 54.2 (#279), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Instruction Following1060—

Long Context Not comparable

Llama2 70b Steerlm Chat: 30.4 (#288), Mistral Nemo: —

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Longer Query998—

Writing & Preference Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.6 (#283), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatMistral Nemo
LMArena Text1098—
LMArena Creative Writing1091—
EQ-Bench Creative Writing—881
LMArena Multi-Turn1058—

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Mistral Nemo?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 26.4 on the Noometry Index.

How many benchmarks do Llama2 70b Steerlm Chat and Mistral Nemo share?

0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper