Model comparison

Longcat Flash Chat vs Qwen1.5-110B

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 34.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Longcat Flash Chat scores higher in 7 categories and Qwen1.5-110B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 38.0.

Side by side

Longcat Flash Chat and Qwen1.5-110B specifications
Longcat Flash ChatQwen1.5-110B
ProviderMeituanAlibaba (Qwen)
Noometry Index42.134.2
Released—2024-04-25
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1920

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen1.5-110B: 33.0 (#264)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Coding14711184
BigCodeBench Instruct—35%
BigCodeBench Complete—44.4%

Reasoning Qwen1.5-110B leads

Longcat Flash Chat: 19.0 (#272), Qwen1.5-110B: 22.7 (#189)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Hard Prompts14401168
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—
ForecastBench—57.7

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Qwen1.5-110B: 33.7 (#201)

Math benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Math14421185

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Qwen1.5-110B: 31.2 (#219)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Expert14541144

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen1.5-110B: 33.6 (#250)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Non-English14041142
LMArena Chinese14651206
LMArena French14561151
LMArena German14081123
LMArena Japanese13731074
LMArena Korean13711044
LMArena Russian13951118
LMArena Spanish14451142

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen1.5-110B: 60.3 (#252)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Instruction Following14111158

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen1.5-110B: 35.1 (#242)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Longer Query14251157

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen1.5-110B: 38.0 (#255)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen1.5-110B
LMArena Text14271175
LMArena Creative Writing13881148
LMArena Multi-Turn14181160

Frequently asked questions

Is Longcat Flash Chat better than Qwen1.5-110B?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 34.2 on the Noometry Index.

Is Longcat Flash Chat or Qwen1.5-110B better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 33.0 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen1.5-110B share?

17 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen1.5-110B has 20.

Related comparisons

Go deeper