Model comparison

Longcat Flash Chat vs Qwen Max

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 34.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Longcat Flash Chat scores higher in 7 categories and Qwen Max in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Longcat Flash Chat leads 39.4 to 22.3.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and Qwen Max specifications
Longcat Flash ChatQwen Max
ProviderMeituanAlibaba (Qwen)
Noometry Index42.134.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1923

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Coding14711288
Aider Polyglot—21.8%

Reasoning Qwen Max leads

Longcat Flash Chat: 19.0 (#272), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Hard Prompts14401269
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Math14421275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Expert14541248
GPQA Diamond—56.1%

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Non-English14041263
LMArena Chinese14651254
LMArena French14561330
LMArena German14081254
LMArena Japanese13731205
LMArena Korean13711142
LMArena Russian13951274
LMArena Spanish14451290

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Instruction Following14111262

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Longer Query14251288
Fiction.LiveBench—66.7%

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen Max
LMArena Text14271282
LMArena Creative Writing13881248
LMArena Multi-Turn14181277

Frequently asked questions

Is Longcat Flash Chat better than Qwen Max?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 34.7 on the Noometry Index.

Is Longcat Flash Chat or Qwen Max better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 30.7 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen Max share?

17 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper