Model comparison

Claude 3 Haiku vs Falcon-180B

Falcon-180B is the stronger model overall, scoring 32.2 to 25.9 on the Noometry Index.

Last verified . 9 shared benchmarks.

Claude 3 Haiku Anthropic

25.9

Rank #340 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Claude 3 Haiku scores higher in 3 categories and Falcon-180B in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Claude 3 Haiku leads 36.0 to 25.2.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Claude 3 Haiku and Falcon-180B specifications
Claude 3 HaikuFalcon-180B
ProviderAnthropicTechnology Innovation Institute
Noometry Index25.932.2
Released2024-03-072023-09-06
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3716

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 3 Haiku: 26.4 (#325), Falcon-180B: —

Coding benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
WeirdML9.8%—
BigCodeBench Instruct39.4%—
LMArena Coding1199—
BigCodeBench Complete50.1%—
CadEval12%—
HumanEval+68.9%—
MBPP+68.8%—

Reasoning Falcon-180B leads

Claude 3 Haiku: 16.3 (#307), Falcon-180B: 19.1 (#269)

Reasoning benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
LMArena Hard Prompts11741007
Epoch Capabilities Index118.35112.13
WinoGrande74.2%87.1%
Kagi LLM Benchmark34.2%—
DTBench50.1%—
LMCA8.8%—
ForecastBench53.2—
HellaSwag—89%
LAMBADA—79.8%
PIQA—84.9%

Math Not comparable

Claude 3 Haiku: 9.8 (#319), Falcon-180B: —

Math benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
OTIS Mock AIME 2024-20251.8%—
LMArena Math1188—
MATH Level 514.9%—
GSM8K—54.4%

Knowledge Not comparable

Claude 3 Haiku: 17.3 (#285), Falcon-180B: —

Knowledge benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
MMLU73.8%70.6%
GPQA Diamond36.3%—
Confabulations34.2%—
LMArena Expert1148—
ARC (AI2) Challenge—67.8%
BoolQ—89%
OpenBookQA—64.2%

Multimodal Not comparable

Claude 3 Haiku: 23.6 (#128), Falcon-180B: —

Multimodal benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
LMArena Vision950—
ScienceQA72%—

Multilingual Claude 3 Haiku leads

Claude 3 Haiku: 36.0 (#243), Falcon-180B: 25.2 (#286)

Multilingual benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
LMArena Non-English11781000
LMArena Chinese1155—
LMArena French1195—
LMArena German1174—
LMArena Japanese1102—
LMArena Korean1109—
LMArena Russian1204—
LMArena Spanish1166—

Instruction Following Claude 3 Haiku leads

Claude 3 Haiku: 61.3 (#247), Falcon-180B: 53.4 (#286)

Instruction Following benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
LMArena Instruction Following11731047

Long Context Not comparable

Claude 3 Haiku: 36.1 (#237), Falcon-180B: —

Long Context benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
LMArena Longer Query1190—

Writing & Preference Too close to call

Claude 3 Haiku: 29.7 (#291), Falcon-180B: 29.1 (#295)

Writing & Preference benchmarks
BenchmarkClaude 3 HaikuFalcon-180B
LMArena Text11951054
LMArena Creative Writing11571089
LMArena Multi-Turn11901013
EQ-Bench Creative Writing717—

Frequently asked questions

Is Claude 3 Haiku better than Falcon-180B?

Falcon-180B is the stronger model overall, scoring 32.2 to 25.9 on the Noometry Index.

How many benchmarks do Claude 3 Haiku and Falcon-180B share?

9 benchmarks have published results for both models. Claude 3 Haiku has 37 scored results on Noometry and Falcon-180B has 16.

Related comparisons

Go deeper