Model comparison

Llama 13b vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 0 categories and Llama2 70b Steerlm Chat in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama2 70b Steerlm Chat leads 31.6 to 13.8.

Side by side

Llama 13b and Llama2 70b Steerlm Chat specifications
Llama 13bLlama2 70b Steerlm Chat
ProviderMetaNVIDIA
Noometry Index24.431.8
Released2023-02-24—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama2 70b Steerlm Chat leads

Llama 13b: 21.4 (#337), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Coding6831025

Reasoning Llama2 70b Steerlm Chat leads

Llama 13b: 14.0 (#329), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Hard Prompts7281047
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama2 70b Steerlm Chat leads

Llama 13b: 26.7 (#256), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Math8381072
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Llama2 70b Steerlm Chat: —

Multimodal benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
ScienceQA43.3%—

Multilingual Llama2 70b Steerlm Chat leads

Llama 13b: 16.6 (#297), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Non-English8191063

Instruction Following Llama2 70b Steerlm Chat leads

Llama 13b: 36.7 (#305), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Instruction Following7811060

Long Context Not comparable

Llama 13b: —, Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Longer Query—998

Writing & Preference Llama2 70b Steerlm Chat leads

Llama 13b: 13.8 (#312), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkLlama 13bLlama2 70b Steerlm Chat
LMArena Text8341098
LMArena Creative Writing7941091
LMArena Multi-Turn7531058

Frequently asked questions

Is Llama 13b better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 24.4 on the Noometry Index.

Is Llama 13b or Llama2 70b Steerlm Chat better for coding?

Llama2 70b Steerlm Chat scores higher on coding benchmarks: 29.9 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Llama2 70b Steerlm Chat share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper