Model comparison

Llama2 70b Steerlm Chat vs Palm 2

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.0 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Palm 2 Google

30.0

Rank #298 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 6 categories and Palm 2 in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Llama2 70b Steerlm Chat leads 28.8 to 20.9.
  • Llama2 70b Steerlm Chat has downloadable open weights; the other is API-only.

Side by side

Llama2 70b Steerlm Chat and Palm 2 specifications
Llama2 70b Steerlm ChatPalm 2
ProviderNVIDIAGoogle
Noometry Index31.830.0
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked910

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama2 70b Steerlm Chat: 29.9 (#300), Palm 2: 29.0 (#311)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Coding1025994

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Palm 2: 19.1 (#271)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Hard Prompts10471005

Math Too close to call

Llama2 70b Steerlm Chat: 31.3 (#226), Palm 2: 30.8 (#231)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Math10721049

Multilingual Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 28.8 (#270), Palm 2: 20.9 (#295)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Non-English1063916
LMArena Chinese—887

Instruction Following Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 54.2 (#279), Palm 2: 51.0 (#296)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Instruction Following10601010

Long Context Too close to call

Llama2 70b Steerlm Chat: 30.4 (#288), Palm 2: 30.8 (#285)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Longer Query9981010

Writing & Preference Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.6 (#283), Palm 2: 25.6 (#304)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatPalm 2
LMArena Text10981027
LMArena Creative Writing1091991
LMArena Multi-Turn1058994

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Palm 2?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.0 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Palm 2 better for coding?

They score almost the same on coding (29.9 vs 29.0); test both on your own repository before choosing.

How many benchmarks do Llama2 70b Steerlm Chat and Palm 2 share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Palm 2 has 10.

Related comparisons

Go deeper