Model comparison

Falcon-180B vs Llama-3.3-70B-Instruct

Falcon-180B is the stronger model overall, scoring 32.2 to 30.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Falcon-180B scores higher in 1 category and Llama-3.3-70B-Instruct in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama-3.3-70B-Instruct leads 47.6 to 29.1.

Side by side

Falcon-180B and Llama-3.3-70B-Instruct specifications
Falcon-180BLlama-3.3-70B-Instruct
ProviderTechnology Innovation InstituteMeta
Noometry Index32.230.6
Released2023-09-062024-12-06
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.32
Results tracked1643

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
SciCode—26%
WeirdML—14.4%
BigCodeBench Instruct—46.9%
LiveBench Coding—36.6%
LMArena Coding—1268
BigCodeBench Complete—57.5%

Agentic & Tool Use Not comparable

Falcon-180B: —, Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
BALROG—23%

Reasoning Falcon-180B leads

Falcon-180B: 19.1 (#269), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
LMArena Hard Prompts10071257
Epoch Capabilities Index112.13127.33
SimpleBench—19.9%
CritPt—0%
LiveBench Reasoning—50.8%
DTBench—59.5%
LiveBench Data Analysis—49.5%
LMCA—17.5%
ForecastBench—58.6
HellaSwag89%—
LAMBADA79.8%—
LiveBench—50.2%
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
OTIS Mock AIME 2024-2025—5.1%
LiveBench Math—42.2%
LMArena Math—1267
MATH Level 5—41.6%
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
MMLU70.6%86.3%
GPQA Diamond—47.4%
Confabulations—22.8%
Vectara Hallucination Rate—4.1%
LMArena Expert—1225
ARC (AI2) Challenge67.8%—
BoolQ89%—
OpenBookQA64.2%—

Multilingual Llama-3.3-70B-Instruct leads

Falcon-180B: 25.2 (#286), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
LMArena Non-English10001236
LMArena Chinese—1217
LMArena French—1281
LMArena German—1251
LMArena Japanese—1150
LMArena Korean—1143
LMArena Russian—1252
LMArena Spanish—1270

Instruction Following Llama-3.3-70B-Instruct leads

Falcon-180B: 53.4 (#286), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
LMArena Instruction Following10471242
LiveBench Instruction Following—82.7%

Long Context Not comparable

Falcon-180B: —, Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
Fiction.LiveBench—33.3%
LMArena Longer Query—1256

Writing & Preference Llama-3.3-70B-Instruct leads

Falcon-180B: 29.1 (#295), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkFalcon-180BLlama-3.3-70B-Instruct
LMArena Text10541274
LMArena Creative Writing10891250
LMArena Multi-Turn10131280
LiveBench Language—39.2%

Frequently asked questions

Is Falcon-180B better than Llama-3.3-70B-Instruct?

Falcon-180B is the stronger model overall, scoring 32.2 to 30.6 on the Noometry Index.

How many benchmarks do Falcon-180B and Llama-3.3-70B-Instruct share?

8 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper