Model comparison

Longcat Flash Chat vs Qwen3 32B

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 39.2 on the Noometry Index.

Last verified . 14 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Longcat Flash Chat scores higher in 5 categories and Qwen3 32B in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 52.9.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 43.9% for Longcat Flash Chat and 54.9% for Qwen3 32B.

Side by side

Longcat Flash Chat and Qwen3 32B specifications
Longcat Flash ChatQwen3 32B
ProviderMeituanAlibaba (Qwen)
Noometry Index42.139.2
Released—2025-04
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked1926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Coding14711358
Aider Polyglot—40%
SciCode—35.4%

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
Berkeley Function Calling Leaderboard—48.7%

Reasoning Qwen3 32B leads

Longcat Flash Chat: 19.0 (#272), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
Kagi LLM Benchmark43.9%54.9%
LMArena Hard Prompts14401334
NYT Connections (extended)17.7%—
CritPt—0.3%
Chess Puzzles—5%
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Too close to call

Longcat Flash Chat: 39.4 (#107), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Math14421399
OTIS Mock AIME 2024-2025—66.9%

Knowledge Too close to call

Longcat Flash Chat: 40.6 (#116), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Expert14541362
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Non-English14041317
LMArena Chinese14651357
LMArena German14081341
LMArena Russian13951311
LMArena French1456—
LMArena Japanese1373—
LMArena Korean1371—
LMArena Spanish1445—

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Instruction Following14111305

Long Context Too close to call

Longcat Flash Chat: 43.5 (#93), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Longer Query14251327
Fiction.LiveBench—74.2%

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatQwen3 32B
LMArena Text14271340
LMArena Creative Writing13881297
LMArena Multi-Turn14181331

Frequently asked questions

Is Longcat Flash Chat better than Qwen3 32B?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 39.2 on the Noometry Index.

Is Longcat Flash Chat or Qwen3 32B better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 37.7 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Qwen3 32B share?

14 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper