Model comparison

Longcat Flash Chat vs Qwen2.5-Max

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 40.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen2.5-Max Alibaba (Qwen)

40.7

Rank #146 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Longcat Flash Chat scores higher in 7 categories and Qwen2.5-Max in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5-Max leads 25.6 to 19.0.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and Qwen2.5-Max specifications
Longcat Flash ChatQwen2.5-Max
ProviderMeituanAlibaba (Qwen)
Noometry Index42.140.7
Released—2025-01-25
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen2.5-Max: 41.8 (#117)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Coding14711359
LiveBench Coding—64.4%

Reasoning Qwen2.5-Max leads

Longcat Flash Chat: 19.0 (#272), Qwen2.5-Max: 25.6 (#147)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Hard Prompts14401360
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—
LiveBench Reasoning—51.4%
LiveBench Data Analysis—67.9%
Epoch Capabilities Index—132.53
LiveBench—62.3%

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Qwen2.5-Max: 36.9 (#162)

Math benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Math14421369
LiveBench Math—58.4%

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Qwen2.5-Max: 35.3 (#186)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Expert14541337
Confabulations—21.8%

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen2.5-Max: 48.1 (#146)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Non-English14041352
LMArena Chinese14651382
LMArena French14561396
LMArena German14081350
LMArena Japanese13731300
LMArena Korean13711304
LMArena Russian13951353
LMArena Spanish14451377

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen2.5-Max: 71.3 (#152)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Instruction Following14111335
LiveBench Instruction Following—75.3%

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen2.5-Max: 41.4 (#142)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Longer Query14251358

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen2.5-Max: 55.4 (#146)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen2.5-Max
LMArena Text14271367
LMArena Creative Writing13881339
LMArena Multi-Turn14181364
Short-Story Creative Writing—72.9%
LiveBench Language—56.3%

Frequently asked questions

Is Longcat Flash Chat better than Qwen2.5-Max?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 40.7 on the Noometry Index.

Is Longcat Flash Chat or Qwen2.5-Max better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 41.8 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen2.5-Max share?

17 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen2.5-Max has 27.

Related comparisons

Go deeper