Model comparison

DeepSeek-V3.1 vs Longcat Flash Chat

DeepSeek-V3.1 and Longcat Flash Chat score almost the same on the Noometry Index (42.8 vs 42.1), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 2 categories and Longcat Flash Chat in 6 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.1 leads 27.9 to 19.0.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 53.2% for DeepSeek-V3.1 and 43.9% for Longcat Flash Chat.

Side by side

DeepSeek-V3.1 and Longcat Flash Chat specifications
DeepSeek-V3.1Longcat Flash Chat
ProviderDeepSeekMeituan
Noometry Index42.842.1
Released2025-08-21—
WeightsOpenOpen
Context window164K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.95—
Results tracked2719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

DeepSeek-V3.1: 40.3 (#144), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Coding14171471
WeirdML38.4%—

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
Kagi LLM Benchmark53.2%43.9%
LMArena Hard Prompts14171440
SimpleBench40%—
NYT Connections (extended)—17.7%
DTBench82.7%—
LMCA24.3%—
Epoch Capabilities Index139.92—
ForecastBench58—

Math Too close to call

DeepSeek-V3.1: 38.9 (#122), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Math14201442

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Expert14051454
Vectara Hallucination Rate5.5%—

Multilingual Too close to call

DeepSeek-V3.1: 51.6 (#106), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Non-English14001404
LMArena Chinese14691465
LMArena French14471456
LMArena German14111408
LMArena Japanese13781373
LMArena Korean13371371
LMArena Russian14051395
LMArena Spanish14311445

Instruction Following Too close to call

DeepSeek-V3.1: 73.9 (#110), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Instruction Following14001411

Long Context Longcat Flash Chat leads

DeepSeek-V3.1: 36.3 (#232), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Longer Query14221425
Fiction.LiveBench52.8%—

Writing & Preference Too close to call

DeepSeek-V3.1: 60.3 (#98), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Longcat Flash Chat
LMArena Text14201427
LMArena Creative Writing14011388
LMArena Multi-Turn14081418
EQ-Bench Creative Writing1436—

Frequently asked questions

Is DeepSeek-V3.1 better than Longcat Flash Chat?

DeepSeek-V3.1 and Longcat Flash Chat score almost the same on the Noometry Index (42.8 vs 42.1), so choose on price, context window or the category you care about most.

Is DeepSeek-V3.1 or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 40.3 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.1 and Longcat Flash Chat share?

18 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper