Model comparison

Falcon-180B vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 32.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Grok 4.1 Fast in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 Fast leads 57.2 to 29.1.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Falcon-180B and Grok 4.1 Fast specifications
Falcon-180BGrok 4.1 Fast
ProviderTechnology Innovation InstitutexAI
Noometry Index32.241.4
Released2023-09-062025-06-27
WeightsOpenProprietary
Context window—128K
Max output—30K
Input $ / M tokens—$0.20
Output $ / M tokens—$0.50
Results tracked1632

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena WebDev—1242
LMArena Coding—1411
ALE-Bench—394.93

Agentic & Tool Use Not comparable

Falcon-180B: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Falcon-180B: 19.1 (#269), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena Hard Prompts10071407
SimpleBench—56%
NYT Connections (extended)—87.4%
DTBench—87.7%
Epoch Capabilities Index112.13—
ForecastBench—61
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
MathArena Final-Answer Competitions—60.9%
ProofBench—4%
LMArena Math—1408
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
Vectara Hallucination Rate—17.8%
LMArena Expert—1399
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multimodal Not comparable

Falcon-180B: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena Vision—1201

Multilingual Grok 4.1 Fast leads

Falcon-180B: 25.2 (#286), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena Non-English10001391
LMArena Chinese—1441
LMArena French—1415
LMArena German—1404
LMArena Japanese—1349
LMArena Korean—1361
LMArena Russian—1387
LMArena Spanish—1413

Instruction Following Grok 4.1 Fast leads

Falcon-180B: 53.4 (#286), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena Instruction Following10471376

Long Context Not comparable

Falcon-180B: —, Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena Longer Query—1390

Writing & Preference Grok 4.1 Fast leads

Falcon-180B: 29.1 (#295), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkFalcon-180BGrok 4.1 Fast
LMArena Text10541408
LMArena Creative Writing10891394
LMArena Multi-Turn10131389
EQ-Bench Creative Writing—1327

Frequently asked questions

Is Falcon-180B better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Grok 4.1 Fast share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper