Model comparison

Llama 3.1-405B vs Longcat Flash Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 30.7 on the Noometry Index.

Last verified . 18 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Llama 3.1-405B scores higher in 0 categories and Longcat Flash Chat in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 38.9.

Side by side

Llama 3.1-405B and Longcat Flash Chat specifications
Llama 3.1-405BLongcat Flash Chat
ProviderMetaMeituan
Noometry Index30.742.1
Released2024-07-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Llama 3.1-405B: 33.1 (#262), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Coding12911471
WeirdML21.4%—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), Longcat Flash Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Longcat Flash Chat leads

Llama 3.1-405B: 16.8 (#300), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
Kagi LLM Benchmark45%43.9%
LMArena Hard Prompts12691440
SimpleBench23%—
NYT Connections (extended)—17.7%
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Longcat Flash Chat leads

Llama 3.1-405B: 18.4 (#290), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Math12811442
OTIS Mock AIME 2024-20259.7%—
Omni-MATH24.9%—
MATH Level 549.8%—

Knowledge Longcat Flash Chat leads

Llama 3.1-405B: 30.4 (#227), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Expert12431454
GPQA Diamond50.9%—
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Longcat Flash Chat leads

Llama 3.1-405B: 40.7 (#214), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Non-English12481404
LMArena Chinese12421465
LMArena French12791456
LMArena German12521408
LMArena Japanese12081373
LMArena Korean11841371
LMArena Russian12651395
LMArena Spanish12601445

Instruction Following Longcat Flash Chat leads

Llama 3.1-405B: 65.9 (#214), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Instruction Following12591411
IFEval81.1%—

Long Context Longcat Flash Chat leads

Llama 3.1-405B: 38.4 (#197), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Longer Query12661425

Writing & Preference Longcat Flash Chat leads

Llama 3.1-405B: 38.9 (#251), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BLongcat Flash Chat
LMArena Text12841427
LMArena Creative Writing12621388
LMArena Multi-Turn12971418
EQ-Bench Creative Writing870—
WildBench78.3%—

Frequently asked questions

Is Llama 3.1-405B better than Longcat Flash Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 30.7 on the Noometry Index.

Is Llama 3.1-405B or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 33.1 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Longcat Flash Chat share?

18 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper