Model comparison

Longcat Flash Chat vs Qwen-14B

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Longcat Flash Chat scores higher in 6 categories and Qwen-14B in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 27.6.

Side by side

Longcat Flash Chat and Qwen-14B specifications
Longcat Flash ChatQwen-14B
ProviderMeituanAlibaba (Qwen)
Noometry Index42.131.4
Released—2023-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Coding14711071

Reasoning Too close to call

Longcat Flash Chat: 19.0 (#272), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Hard Prompts14401027
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—
BIG-Bench Hard—55%
Epoch Capabilities Index—113.03
LAMBADA—71.1%
PIQA—79.9%

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Math14421068
GSM8K—61.3%

Knowledge Not comparable

Longcat Flash Chat: 40.6 (#116), Qwen-14B: —

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Expert1454—
ARC (AI2) Challenge—84.4%
BoolQ—86.2%
MMLU—66.3%

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Non-English14041041
LMArena Chinese14651077
LMArena French1456—
LMArena German1408—
LMArena Japanese1373—
LMArena Korean1371—
LMArena Russian1395—
LMArena Spanish1445—

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Instruction Following14111031

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Longer Query14251028

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen-14B
LMArena Text14271051
LMArena Creative Writing13881028
LMArena Multi-Turn14181022

Frequently asked questions

Is Longcat Flash Chat better than Qwen-14B?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 31.4 on the Noometry Index.

Is Longcat Flash Chat or Qwen-14B better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 31.2 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen-14B share?

10 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper