Model comparison

Longcat Flash Chat vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 42.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Longcat Flash Chat scores higher in 3 categories and Qwen3 Max in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 40.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 43.9% for Longcat Flash Chat and 72.5% for Qwen3 Max.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and Qwen3 Max specifications
Longcat Flash ChatQwen3 Max
ProviderMeituanAlibaba (Qwen)
Noometry Index42.143.7
Released—2025-09-23
WeightsOpenProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.20
Output $ / M tokens—$6
Results tracked1933

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Longcat Flash Chat: 43.5 (#87), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Coding14711456
ALE-Bench—370.45

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
Vending-Bench 2—71.56

Reasoning Qwen3 Max leads

Longcat Flash Chat: 19.0 (#272), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
Kagi LLM Benchmark43.9%72.5%
NYT Connections (extended)17.7%30.1%
LMArena Hard Prompts14401448
Chess Puzzles—4%
Mystery Game Puzzles—5%
DTBench—82.1%
LMCA—28.3%
Epoch Capabilities Index—142.38

Math Too close to call

Longcat Flash Chat: 39.4 (#107), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Math14421446
FrontierMath (Tiers 1-3)—18.9%
OTIS Mock AIME 2024-2025—73.3%
MATH Level 5—97.1%

Knowledge Qwen3 Max leads

Longcat Flash Chat: 40.6 (#116), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Expert14541455
GPQA Diamond—72.6%
SimpleQA Verified—48.7%

Multilingual Qwen3 Max leads

Longcat Flash Chat: 51.9 (#101), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Non-English14041429
LMArena Chinese14651478
LMArena French14561449
LMArena German14081463
LMArena Japanese13731397
LMArena Korean13711399
LMArena Russian13951428
LMArena Spanish14451462

Instruction Following Too close to call

Longcat Flash Chat: 74.4 (#96), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Instruction Following14111419

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Longer Query14251438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Longcat Flash Chat: 61.0 (#91), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen3 Max
LMArena Text14271439
LMArena Creative Writing13881402
LMArena Multi-Turn14181446

Frequently asked questions

Is Longcat Flash Chat better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 42.1 on the Noometry Index.

Is Longcat Flash Chat or Qwen3 Max better for coding?

They score almost the same on coding (43.5 vs 43.0); test both on your own repository before choosing.

How many benchmarks do Longcat Flash Chat and Qwen3 Max share?

19 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper