Model comparison

Longcat Flash Chat vs Qwen1.5 4b Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 28.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Longcat Flash Chat scores higher in 8 categories and Qwen1.5 4b Chat in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 23.8.

Side by side

Longcat Flash Chat and Qwen1.5 4b Chat specifications
Longcat Flash ChatQwen1.5 4b Chat
ProviderMeituanAlibaba (Qwen)
Noometry Index42.128.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1913

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Coding1471999

Reasoning Too close to call

Longcat Flash Chat: 19.0 (#272), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Hard Prompts1440976
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Math14421026

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Expert1454980

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Non-English1404979
LMArena Chinese14651024
LMArena German1408902
LMArena Russian1395952
LMArena French1456—
LMArena Japanese1373—
LMArena Korean1371—
LMArena Spanish1445—

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Instruction Following1411978

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Longer Query1425988

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen1.5 4b Chat
LMArena Text1427997
LMArena Creative Writing1388969
LMArena Multi-Turn1418977

Frequently asked questions

Is Longcat Flash Chat better than Qwen1.5 4b Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 28.8 on the Noometry Index.

Is Longcat Flash Chat or Qwen1.5 4b Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 29.1 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper