Model comparison

Falcon-180B vs Llama 3.1 Nemotron 70b Instruct

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 32.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Llama 3.1 Nemotron 70b Instruct in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.1 Nemotron 70b Instruct leads 48.4 to 29.1.

Side by side

Falcon-180B and Llama 3.1 Nemotron 70b Instruct specifications
Falcon-180BLlama 3.1 Nemotron 70b Instruct
ProviderTechnology Innovation InstituteNVIDIA
Noometry Index32.237.6
Released2023-09-062024-12-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1614

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 70b Instruct: 35.9 (#216)

Coding benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
BigCodeBench Instruct—38.7%
LMArena Coding—1272
BigCodeBench Complete—48.2%

Reasoning Llama 3.1 Nemotron 70b Instruct leads

Falcon-180B: 19.1 (#269), Llama 3.1 Nemotron 70b Instruct: 25.0 (#152)

Reasoning benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Hard Prompts10071266
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 70b Instruct: 35.5 (#182)

Math benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Math—1271
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 70b Instruct: 34.1 (#199)

Knowledge benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Expert—1242
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Llama 3.1 Nemotron 70b Instruct leads

Falcon-180B: 25.2 (#286), Llama 3.1 Nemotron 70b Instruct: 40.5 (#217)

Multilingual benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Non-English10001245
LMArena Chinese—1263
LMArena Russian—1227

Instruction Following Llama 3.1 Nemotron 70b Instruct leads

Falcon-180B: 53.4 (#286), Llama 3.1 Nemotron 70b Instruct: 65.9 (#213)

Instruction Following benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Instruction Following10471252

Long Context Not comparable

Falcon-180B: —, Llama 3.1 Nemotron 70b Instruct: 37.6 (#215)

Long Context benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Longer Query—1238

Writing & Preference Llama 3.1 Nemotron 70b Instruct leads

Falcon-180B: 29.1 (#295), Llama 3.1 Nemotron 70b Instruct: 48.4 (#203)

Writing & Preference benchmarks
BenchmarkFalcon-180BLlama 3.1 Nemotron 70b Instruct
LMArena Text10541283
LMArena Creative Writing10891269
LMArena Multi-Turn10131275

Frequently asked questions

Is Falcon-180B better than Llama 3.1 Nemotron 70b Instruct?

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Llama 3.1 Nemotron 70b Instruct share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Llama 3.1 Nemotron 70b Instruct has 14.

Related comparisons

Go deeper