Model comparison

Falcon-180B vs Llama 2-7B

Falcon-180B is the stronger model overall, scoring 32.2 to 29.1 on the Noometry Index.

Last verified . 16 shared benchmarks.

Llama 2-7B Meta

29.1

Rank #317 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Falcon-180B scores higher in 4 categories and Llama 2-7B in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Falcon-180B leads 19.1 to 15.7.

Side by side

Falcon-180B and Llama 2-7B specifications
Falcon-180BLlama 2-7B
ProviderTechnology Innovation InstituteMeta
Noometry Index32.229.1
Released2023-09-062023-07-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1629

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Llama 2-7B: 29.2 (#307)

Coding benchmarks
BenchmarkFalcon-180BLlama 2-7B
LMArena Coding—1002

Reasoning Falcon-180B leads

Falcon-180B: 19.1 (#269), Llama 2-7B: 15.7 (#312)

Reasoning benchmarks
BenchmarkFalcon-180BLlama 2-7B
LMArena Hard Prompts10071009
Epoch Capabilities Index112.1399.06
HellaSwag89%77.2%
LAMBADA79.8%73.3%
PIQA84.9%78.8%
WinoGrande87.1%69.2%
Chess Puzzles—0%
BIG-Bench Hard—39.2%

Math Not comparable

Falcon-180B: —, Llama 2-7B: 30.7 (#233)

Math benchmarks
BenchmarkFalcon-180BLlama 2-7B
GSM8K54.4%16.7%
LMArena Math—1042

Knowledge Not comparable

Falcon-180B: —, Llama 2-7B: 28.2 (#248)

Knowledge benchmarks
BenchmarkFalcon-180BLlama 2-7B
ARC (AI2) Challenge67.8%45.9%
BoolQ89%77.9%
MMLU70.6%45.8%
OpenBookQA64.2%58.6%
LMArena Expert—1036
TriviaQA—73.7%

Multimodal Not comparable

Falcon-180B: —, Llama 2-7B: —

Multimodal benchmarks
BenchmarkFalcon-180BLlama 2-7B
ScienceQA—43.1%

Multilingual Falcon-180B leads

Falcon-180B: 25.2 (#286), Llama 2-7B: 23.8 (#293)

Multilingual benchmarks
BenchmarkFalcon-180BLlama 2-7B
LMArena Non-English1000973
LMArena Chinese—973
LMArena French—970
LMArena German—978
LMArena Russian—995
LMArena Spanish—1007

Instruction Following Falcon-180B leads

Falcon-180B: 53.4 (#286), Llama 2-7B: 50.8 (#298)

Instruction Following benchmarks
BenchmarkFalcon-180BLlama 2-7B
LMArena Instruction Following10471006

Long Context Not comparable

Falcon-180B: —, Llama 2-7B: 30.4 (#287)

Long Context benchmarks
BenchmarkFalcon-180BLlama 2-7B
LMArena Longer Query—999

Writing & Preference Falcon-180B leads

Falcon-180B: 29.1 (#295), Llama 2-7B: 28.0 (#298)

Writing & Preference benchmarks
BenchmarkFalcon-180BLlama 2-7B
LMArena Text10541053
LMArena Creative Writing10891033
LMArena Multi-Turn10131029

Frequently asked questions

Is Falcon-180B better than Llama 2-7B?

Falcon-180B is the stronger model overall, scoring 32.2 to 29.1 on the Noometry Index.

How many benchmarks do Falcon-180B and Llama 2-7B share?

16 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Llama 2-7B has 29.

Related comparisons

Go deeper