Model comparison

Falcon-180B vs Llama2 70b Steerlm Chat

Falcon-180B and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (32.2 vs 31.8), so choose on price, context window or the category you care about most.

Last verified . 6 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Llama2 70b Steerlm Chat in 4 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Llama2 70b Steerlm Chat leads 28.8 to 25.2.

Side by side

Falcon-180B and Llama2 70b Steerlm Chat specifications
Falcon-180BLlama2 70b Steerlm Chat
ProviderTechnology Innovation InstituteNVIDIA
Noometry Index32.231.8
Released2023-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked169

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Coding—1025

Reasoning Too close to call

Falcon-180B: 19.1 (#269), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Hard Prompts10071047
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Math—1072
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Llama2 70b Steerlm Chat leads

Falcon-180B: 25.2 (#286), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Non-English10001063

Instruction Following Too close to call

Falcon-180B: 53.4 (#286), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Instruction Following10471060

Long Context Not comparable

Falcon-180B: —, Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Longer Query—998

Writing & Preference Llama2 70b Steerlm Chat leads

Falcon-180B: 29.1 (#295), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkFalcon-180BLlama2 70b Steerlm Chat
LMArena Text10541098
LMArena Creative Writing10891091
LMArena Multi-Turn10131058

Frequently asked questions

Is Falcon-180B better than Llama2 70b Steerlm Chat?

Falcon-180B and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (32.2 vs 31.8), so choose on price, context window or the category you care about most.

How many benchmarks do Falcon-180B and Llama2 70b Steerlm Chat share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper