Model comparison

Falcon-180B vs o3-pro

o3-pro is the stronger model overall, scoring 42.9 to 32.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • They share 1 benchmark with published results for both. Falcon-180B scores higher in 0 categories and o3-pro in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where o3-pro leads 57.1 to 29.1.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Falcon-180B and o3-pro specifications
Falcon-180Bo3-pro
ProviderTechnology Innovation InstituteOpenAI
Noometry Index32.242.9
Released2023-09-062025-06-10
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$20
Output $ / M tokens—$80
Results tracked1612

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkFalcon-180Bo3-pro
Aider Polyglot—84.9%
WeirdML—58.2%

Reasoning o3-pro leads

Falcon-180B: 19.1 (#269), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkFalcon-180Bo3-pro
Epoch Capabilities Index112.13147.42
ARC-AGI-2—4.9%
Kagi LLM Benchmark—72.1%
ARC-AGI-1—59.3%
LMArena Hard Prompts1007—
DTBench—86.9%
LMCA—38.5%
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, o3-pro: —

Math benchmarks
BenchmarkFalcon-180Bo3-pro
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkFalcon-180Bo3-pro
Confabulations—14.2%
Vectara Hallucination Rate—23.3%
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Not comparable

Falcon-180B: 25.2 (#286), o3-pro: —

Multilingual benchmarks
BenchmarkFalcon-180Bo3-pro
LMArena Non-English1000—

Instruction Following Not comparable

Falcon-180B: 53.4 (#286), o3-pro: —

Instruction Following benchmarks
BenchmarkFalcon-180Bo3-pro
LMArena Instruction Following1047—

Long Context Not comparable

Falcon-180B: —, o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkFalcon-180Bo3-pro
Fiction.LiveBench—97.2%

Writing & Preference o3-pro leads

Falcon-180B: 29.1 (#295), o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkFalcon-180Bo3-pro
LMArena Text1054—
LMArena Creative Writing1089—
Short-Story Creative Writing—84.4%
LMArena Multi-Turn1013—

Frequently asked questions

Is Falcon-180B better than o3-pro?

o3-pro is the stronger model overall, scoring 42.9 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and o3-pro share?

1 benchmark has published results for both models. Falcon-180B has 16 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper