Model comparison

Longcat Flash Chat vs Qwen2-72B

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 30.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen2-72B Alibaba (Qwen)

30.0

Rank #300 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Longcat Flash Chat scores higher in 7 categories and Qwen2-72B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 40.8.

Side by side

Longcat Flash Chat and Qwen2-72B specifications
Longcat Flash ChatQwen2-72B
ProviderMeituanAlibaba (Qwen)
Noometry Index42.130.0
Released—2024-06-07
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen2-72B: 29.1 (#310)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Coding14711196
WeirdML—11.3%
BigCodeBench Instruct—38.5%
BigCodeBench Complete—54%

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Qwen2-72B: 17.0 (#146)

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
TheAgentCompany—1.1%
METR Time Horizons—29.9%

Reasoning Qwen2-72B leads

Longcat Flash Chat: 19.0 (#272), Qwen2-72B: 23.2 (#181)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Hard Prompts14401191
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—
Epoch Capabilities Index—125.28

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Qwen2-72B: 30.2 (#236)

Math benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Math14421235
MATH Level 5—39.1%

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Qwen2-72B: 21.2 (#275)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Expert14541171
GPQA Diamond—40.8%
MMLU—82.4%

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen2-72B: 35.9 (#244)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Non-English14041176
LMArena Chinese14651240
LMArena French14561170
LMArena German14081151
LMArena Japanese13731111
LMArena Korean13711083
LMArena Russian13951169
LMArena Spanish14451169

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen2-72B: 61.7 (#241)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Instruction Following14111181

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen2-72B: 36.1 (#235)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Longer Query14251192

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen2-72B: 40.8 (#241)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen2-72B
LMArena Text14271203
LMArena Creative Writing13881181
LMArena Multi-Turn14181196

Frequently asked questions

Is Longcat Flash Chat better than Qwen2-72B?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 30.0 on the Noometry Index.

Is Longcat Flash Chat or Qwen2-72B better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 29.1 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen2-72B share?

17 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen2-72B has 26.

Related comparisons

Go deeper