Model comparison

Llama 3-8B vs Longcat Flash Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 25.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 3-8B scores higher in 0 categories and Longcat Flash Chat in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Longcat Flash Chat leads 40.6 to 7.8.

Side by side

Llama 3-8B and Longcat Flash Chat specifications
Llama 3-8BLongcat Flash Chat
ProviderMetaMeituan
Noometry Index25.542.1
Released2024-04-18—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Llama 3-8B: 31.0 (#289), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Coding11521471
BigCodeBench Instruct31.9%—
BigCodeBench Complete36.9%—
HumanEval+56.7%—
MBPP+54.8%—

Reasoning Longcat Flash Chat leads

Llama 3-8B: 14.3 (#326), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Hard Prompts11331440
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%
Chess Puzzles0%—
DTBench43.9%—
Adversarial NLI57.3%—
Epoch Capabilities Index116.45—
ForecastBench58.6—
WinoGrande75.7%—

Math Longcat Flash Chat leads

Llama 3-8B: 8.8 (#323), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Math11511442
OTIS Mock AIME 2024-20251.9%—
MATH Level 56.1%—

Knowledge Longcat Flash Chat leads

Llama 3-8B: 7.8 (#308), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Expert11131454
GPQA Diamond26.1%—
ARC (AI2) Challenge82.8%—
MMLU68.8%—
OpenBookQA82.6%—
TriviaQA67.7%—

Multilingual Longcat Flash Chat leads

Llama 3-8B: 30.8 (#261), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Non-English10981404
LMArena Chinese10761465
LMArena French11591456
LMArena German11041408
LMArena Japanese9671373
LMArena Korean10041371
LMArena Russian11091395
LMArena Spanish11731445

Instruction Following Longcat Flash Chat leads

Llama 3-8B: 58.4 (#260), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Instruction Following11271411

Long Context Longcat Flash Chat leads

Llama 3-8B: 34.2 (#251), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Longer Query11281425

Writing & Preference Longcat Flash Chat leads

Llama 3-8B: 37.5 (#256), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkLlama 3-8BLongcat Flash Chat
LMArena Text11661427
LMArena Creative Writing11501388
LMArena Multi-Turn11521418

Frequently asked questions

Is Llama 3-8B better than Longcat Flash Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 25.5 on the Noometry Index.

Is Llama 3-8B or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 31.0 in the Noometry coding category.

How many benchmarks do Llama 3-8B and Longcat Flash Chat share?

17 benchmarks have published results for both models. Llama 3-8B has 34 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper