Model comparison

Longcat Flash Chat vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 42.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Longcat Flash Chat scores higher in 1 category and Qwen3.6 Plus in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Plus leads 56.1 to 40.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 17.7% for Longcat Flash Chat and 60.3% for Qwen3.6 Plus.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and Qwen3.6 Plus specifications
Longcat Flash ChatQwen3.6 Plus
ProviderMeituanAlibaba (Qwen)
Noometry Index42.147.5
Released—2026-03-31
WeightsOpenProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.50
Output $ / M tokens—$3
Results tracked1937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Coding14711467
SWE-bench Verified—57.9%
LMArena WebDev—1461
SciCode—40.7%
ALE-Bench—670.15

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
Vending-Bench 2—5,115

Reasoning Qwen3.6 Plus leads

Longcat Flash Chat: 19.0 (#272), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
NYT Connections (extended)17.7%60.3%
LMArena Hard Prompts14401449
Kagi LLM Benchmark43.9%—
CritPt—2.9%
Chess Puzzles—17%
Thematic Generalization—59.5%
Mystery Game Puzzles—12%
DTBench—81.9%
LMCA—33.1%
Epoch Capabilities Index—147.65

Math Qwen3.6 Plus leads

Longcat Flash Chat: 39.4 (#107), Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Math14421450
FrontierMath (Tiers 1-3)—38.2%
OTIS Mock AIME 2024-2025—93.3%
FrontierMath (Feb 2025 set)—26.2%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Qwen3.6 Plus leads

Longcat Flash Chat: 40.6 (#116), Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Expert14541454
GPQA Diamond—88.4%
SimpleQA Verified—44.1%

Multilingual Qwen3.6 Plus leads

Longcat Flash Chat: 51.9 (#101), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Non-English14041424
LMArena Chinese14651477
LMArena French14561455
LMArena German14081452
LMArena Japanese13731389
LMArena Korean13711379
LMArena Russian13951434
LMArena Spanish14451432

Instruction Following Too close to call

Longcat Flash Chat: 74.4 (#96), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Instruction Following14111425

Long Context Qwen3.6 Plus leads

Longcat Flash Chat: 43.5 (#93), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Longer Query14251439
CL-bench—20.3%

Writing & Preference Qwen3.6 Plus leads

Longcat Flash Chat: 61.0 (#91), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Plus
LMArena Text14271437
LMArena Creative Writing13881404
LMArena Multi-Turn14181438

Frequently asked questions

Is Longcat Flash Chat better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 42.1 on the Noometry Index.

Is Longcat Flash Chat or Qwen3.6 Plus better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 40.8 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen3.6 Plus share?

18 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper