Model comparison

Falcon-180B vs Gemini 1.0 Pro

Falcon-180B is the stronger model overall, scoring 32.2 to 27.3 on the Noometry Index.

Last verified . 8 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Falcon-180B scores higher in 1 category and Gemini 1.0 Pro in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Gemini 1.0 Pro leads 33.4 to 25.2.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Falcon-180B and Gemini 1.0 Pro specifications
Falcon-180BGemini 1.0 Pro
ProviderTechnology Innovation InstituteGoogle
Noometry Index32.227.3
Released2023-09-062023-12-13
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1624

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Gemini 1.0 Pro: 32.2 (#275)

Coding benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
LMArena Coding—1108
HumanEval+—55.5%
MBPP+—61.4%

Reasoning Falcon-180B leads

Falcon-180B: 19.1 (#269), Gemini 1.0 Pro: 17.1 (#296)

Reasoning benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
LMArena Hard Prompts10071109
Epoch Capabilities Index112.13117.04
DTBench—45.9%
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Gemini 1.0 Pro: 9.3 (#321)

Math benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
OTIS Mock AIME 2024-2025—1.1%
LMArena Math—1132
MATH Level 5—11.2%
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Gemini 1.0 Pro: 15.6 (#291)

Knowledge benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
MMLU70.6%70%
GPQA Diamond—34%
LMArena Expert—1059
ARC (AI2) Challenge67.8%—
BoolQ89%—
OpenBookQA64.2%—

Multilingual Gemini 1.0 Pro leads

Falcon-180B: 25.2 (#286), Gemini 1.0 Pro: 33.4 (#252)

Multilingual benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
LMArena Non-English10001138
LMArena Chinese—1124
LMArena French—1145
LMArena German—1125
LMArena Japanese—1023
LMArena Russian—1186
LMArena Spanish—1119

Instruction Following Gemini 1.0 Pro leads

Falcon-180B: 53.4 (#286), Gemini 1.0 Pro: 57.6 (#267)

Instruction Following benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
LMArena Instruction Following10471114

Long Context Not comparable

Falcon-180B: —, Gemini 1.0 Pro: 34.3 (#249)

Long Context benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
LMArena Longer Query—1132

Writing & Preference Gemini 1.0 Pro leads

Falcon-180B: 29.1 (#295), Gemini 1.0 Pro: 36.0 (#264)

Writing & Preference benchmarks
BenchmarkFalcon-180BGemini 1.0 Pro
LMArena Text10541149
LMArena Creative Writing10891131
LMArena Multi-Turn10131139

Frequently asked questions

Is Falcon-180B better than Gemini 1.0 Pro?

Falcon-180B is the stronger model overall, scoring 32.2 to 27.3 on the Noometry Index.

How many benchmarks do Falcon-180B and Gemini 1.0 Pro share?

8 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Gemini 1.0 Pro has 24.

Related comparisons

Go deeper