Model comparison

Falcon-180B vs Gemini 1.5 Flash (May 2024)

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 32.2 on the Noometry Index.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Gemini 1.5 Flash (May 2024) in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash (May 2024) leads 48.7 to 29.1.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Falcon-180B and Gemini 1.5 Flash (May 2024) specifications
Falcon-180BGemini 1.5 Flash (May 2024)
ProviderTechnology Innovation InstituteGoogle
Noometry Index32.233.2
Released2023-09-062024-05-14
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1642

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
WeirdML—24.9%
BigCodeBench Instruct—43.5%
LMArena Coding—1261
BigCodeBench Complete—55.1%
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Not comparable

Falcon-180B: —, Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
BALROG—14.6%

Reasoning Gemini 1.5 Flash (May 2024) leads

Falcon-180B: 19.1 (#269), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
LMArena Hard Prompts10071257
Epoch Capabilities Index112.13129.36
PIQA84.9%87.5%
DTBench—53.8%
ForecastBench—53.9
HellaSwag89%—
LAMBADA79.8%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
GSM8K54.4%82.4%
OTIS Mock AIME 2024-2025—16.3%
Omni-MATH—30.4%
LMArena Math—1269
MATH Level 5—61.9%
FrontierMath (Feb 2025 set)—0%

Knowledge Not comparable

Falcon-180B: —, Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
BoolQ89%85.8%
MMLU70.6%77.9%
GPQA Diamond—47.3%
MMLU-Pro—67.8%
GPQA (HELM)—43.7%
LMArena Expert—1233
ARC (AI2) Challenge67.8%—
OpenBookQA64.2%—

Multimodal Not comparable

Falcon-180B: —, Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
LMArena Vision—1141
Video-MME—70.3%
GeoBench—76%

Multilingual Gemini 1.5 Flash (May 2024) leads

Falcon-180B: 25.2 (#286), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
LMArena Non-English10001278
LMArena Chinese—1295
LMArena French—1258
LMArena German—1262
LMArena Japanese—1252
LMArena Korean—1221
LMArena Russian—1288
LMArena Spanish—1243

Instruction Following Gemini 1.5 Flash (May 2024) leads

Falcon-180B: 53.4 (#286), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
LMArena Instruction Following10471258
IFEval—83.1%

Long Context Not comparable

Falcon-180B: —, Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
LMArena Longer Query—1284

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Falcon-180B: 29.1 (#295), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkFalcon-180BGemini 1.5 Flash (May 2024)
LMArena Text10541287
LMArena Creative Writing10891285
LMArena Multi-Turn10131253
WildBench—79.2%

Frequently asked questions

Is Falcon-180B better than Gemini 1.5 Flash (May 2024)?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Gemini 1.5 Flash (May 2024) share?

11 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper