Model comparison

Falcon-180B vs Longcat Flash Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 32.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 1 category and Longcat Flash Chat in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 29.1.

Side by side

Falcon-180B and Longcat Flash Chat specifications
Falcon-180BLongcat Flash Chat
ProviderTechnology Innovation InstituteMeituan
Noometry Index32.242.1
Released2023-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Coding—1471

Reasoning Too close to call

Falcon-180B: 19.1 (#269), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Hard Prompts10071440
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Math—1442
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Expert—1454
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Longcat Flash Chat leads

Falcon-180B: 25.2 (#286), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Non-English10001404
LMArena Chinese—1465
LMArena French—1456
LMArena German—1408
LMArena Japanese—1373
LMArena Korean—1371
LMArena Russian—1395
LMArena Spanish—1445

Instruction Following Longcat Flash Chat leads

Falcon-180B: 53.4 (#286), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Instruction Following10471411

Long Context Not comparable

Falcon-180B: —, Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Longer Query—1425

Writing & Preference Longcat Flash Chat leads

Falcon-180B: 29.1 (#295), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkFalcon-180BLongcat Flash Chat
LMArena Text10541427
LMArena Creative Writing10891388
LMArena Multi-Turn10131418

Frequently asked questions

Is Falcon-180B better than Longcat Flash Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Longcat Flash Chat share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper