Model comparison

Falcon-180B vs Granite 3.0 8b Instruct

Falcon-180B and Granite 3.0 8b Instruct score almost the same on the Noometry Index (32.2 vs 31.6), so choose on price, context window or the category you care about most.

Last verified . 6 shared benchmarks.

Granite 3.0 8b Instruct IBM

31.6

Rank #270 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Granite 3.0 8b Instruct in 4 categories; 4 gaps are clear of the uncertainty.

Side by side

Falcon-180B and Granite 3.0 8b Instruct specifications
Falcon-180BGranite 3.0 8b Instruct
ProviderTechnology Innovation InstituteIBM
Noometry Index32.231.6
Released2023-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1614

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Granite 3.0 8b Instruct: 29.7 (#301)

Coding benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
BigCodeBench Instruct—29.3%
LMArena Coding—1112
BigCodeBench Complete—35.4%

Reasoning Granite 3.0 8b Instruct leads

Falcon-180B: 19.1 (#269), Granite 3.0 8b Instruct: 20.9 (#230)

Reasoning benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Hard Prompts10071092
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Granite 3.0 8b Instruct: 32.8 (#210)

Math benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Math—1143
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Granite 3.0 8b Instruct: 29.6 (#236)

Knowledge benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Expert—1087
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Granite 3.0 8b Instruct leads

Falcon-180B: 25.2 (#286), Granite 3.0 8b Instruct: 27.2 (#276)

Multilingual benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Non-English10001037
LMArena Chinese—1063
LMArena Russian—1060

Instruction Following Granite 3.0 8b Instruct leads

Falcon-180B: 53.4 (#286), Granite 3.0 8b Instruct: 56.0 (#276)

Instruction Following benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Instruction Following10471088

Long Context Not comparable

Falcon-180B: —, Granite 3.0 8b Instruct: 34.0 (#252)

Long Context benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Longer Query—1122

Writing & Preference Granite 3.0 8b Instruct leads

Falcon-180B: 29.1 (#295), Granite 3.0 8b Instruct: 31.1 (#285)

Writing & Preference benchmarks
BenchmarkFalcon-180BGranite 3.0 8b Instruct
LMArena Text10541096
LMArena Creative Writing10891071
LMArena Multi-Turn10131063

Frequently asked questions

Is Falcon-180B better than Granite 3.0 8b Instruct?

Falcon-180B and Granite 3.0 8b Instruct score almost the same on the Noometry Index (32.2 vs 31.6), so choose on price, context window or the category you care about most.

How many benchmarks do Falcon-180B and Granite 3.0 8b Instruct share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Granite 3.0 8b Instruct has 14.

Related comparisons

Go deeper