Model comparison

Falcon-180B vs Llama 3.1 Nemotron 51b Instruct

Llama 3.1 Nemotron 51b Instruct is the stronger model overall, scoring 35.9 to 32.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Llama 3.1 Nemotron 51b Instruct in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.1 Nemotron 51b Instruct leads 43.4 to 29.1.

Side by side

Falcon-180B and Llama 3.1 Nemotron 51b Instruct specifications
Falcon-180BLlama 3.1 Nemotron 51b Instruct
ProviderTechnology Innovation InstituteNVIDIA
Noometry Index32.235.9
Released2023-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1612

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 51b Instruct: 35.6 (#222)

Coding benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Coding—1223

Reasoning Llama 3.1 Nemotron 51b Instruct leads

Falcon-180B: 19.1 (#269), Llama 3.1 Nemotron 51b Instruct: 23.5 (#177)

Reasoning benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Hard Prompts10071203
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 51b Instruct: 34.6 (#193)

Math benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Math—1230
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 51b Instruct: 31.9 (#218)

Knowledge benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Expert—1167
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Llama 3.1 Nemotron 51b Instruct leads

Falcon-180B: 25.2 (#286), Llama 3.1 Nemotron 51b Instruct: 36.1 (#241)

Multilingual benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Non-English10001181
LMArena Chinese—1180
LMArena Russian—1187

Instruction Following Llama 3.1 Nemotron 51b Instruct leads

Falcon-180B: 53.4 (#286), Llama 3.1 Nemotron 51b Instruct: 62.9 (#233)

Instruction Following benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Instruction Following10471201

Long Context Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 51b Instruct: 36.5 (#230)

Long Context benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Longer Query—1205

Writing & Preference Llama 3.1 Nemotron 51b Instruct leads

Falcon-180B: 29.1 (#295), Llama 3.1 Nemotron 51b Instruct: 43.4 (#229)

Writing & Preference benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 51b Instruct
LMArena Text10541228
LMArena Creative Writing10891213
LMArena Multi-Turn10131227

Frequently asked questions

Is Falcon-180B better than Llama 3.1 Nemotron 51b Instruct?

Llama 3.1 Nemotron 51b Instruct is the stronger model overall, scoring 35.9 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Llama 3.1 Nemotron 51b Instruct share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Llama 3.1 Nemotron 51b Instruct has 12.

Related comparisons

Go deeper