Model comparison

Longcat Flash Chat vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 42.1 on the Noometry Index.

Last verified . 15 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Longcat Flash Chat scores higher in 0 categories and Qwen3.6 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Max Preview leads 41.7 to 19.0.
  • The biggest single-benchmark swing is NYT Connections (extended): 17.7% for Longcat Flash Chat and 74.1% for Qwen3.6 Max Preview.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and Qwen3.6 Max Preview specifications
Longcat Flash ChatQwen3.6 Max Preview
ProviderMeituanAlibaba (Qwen)
Noometry Index42.151.5
Released—2026-04-20
WeightsOpenProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.30
Output $ / M tokens—$7.80
Results tracked1929

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Longcat Flash Chat: 43.5 (#87), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Coding14711471
SWE-bench Verified—76.7%
LMArena WebDev—1482

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Longcat Flash Chat: 19.0 (#272), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
NYT Connections (extended)17.7%74.1%
LMArena Hard Prompts14401457
SimpleBench—63%
Kagi LLM Benchmark43.9%—
Chess Puzzles—20%
Mystery Game Puzzles—19%
DTBench—87.2%
LMCA—42.5%
Epoch Capabilities Index—149.24

Math Qwen3.6 Max Preview leads

Longcat Flash Chat: 39.4 (#107), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Math14421465
OTIS Mock AIME 2024-2025—91.1%
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Qwen3.6 Max Preview leads

Longcat Flash Chat: 40.6 (#116), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Expert14541478
GPQA Diamond—87.4%
SimpleQA Verified—52%

Multilingual Qwen3.6 Max Preview leads

Longcat Flash Chat: 51.9 (#101), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Non-English14041437
LMArena Chinese14651487
LMArena French14561449
LMArena Russian13951445
LMArena Spanish14451454
LMArena German1408—
LMArena Japanese1373—
LMArena Korean1371—

Instruction Following Qwen3.6 Max Preview leads

Longcat Flash Chat: 74.4 (#96), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Instruction Following14111438

Long Context Qwen3.6 Max Preview leads

Longcat Flash Chat: 43.5 (#93), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Longer Query14251457

Writing & Preference Qwen3.6 Max Preview leads

Longcat Flash Chat: 61.0 (#91), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen3.6 Max Preview
LMArena Text14271447
LMArena Creative Writing13881435
LMArena Multi-Turn14181456

Frequently asked questions

Is Longcat Flash Chat better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 42.1 on the Noometry Index.

Is Longcat Flash Chat or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 43.5 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen3.6 Max Preview share?

15 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper