Model comparison

Longcat Flash Chat vs Mistral Large 3

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 39.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Longcat Flash Chat scores higher in 7 categories and Mistral Large 3 in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Longcat Flash Chat leads 43.5 to 34.4.
  • The biggest single-benchmark swing is NYT Connections (extended): 17.7% for Longcat Flash Chat and 7.5% for Mistral Large 3.

Side by side

Longcat Flash Chat and Mistral Large 3 specifications
Longcat Flash ChatMistral Large 3
ProviderMeituanMistral AI
Noometry Index42.139.1
Released—2025-12-02
WeightsOpenOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked1924

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Coding14711448
LMArena WebDev—1230

Reasoning Longcat Flash Chat leads

Longcat Flash Chat: 19.0 (#272), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
Kagi LLM Benchmark43.9%50.9%
NYT Connections (extended)17.7%7.5%
LMArena Hard Prompts14401429
Thematic Generalization—23%

Math Too close to call

Longcat Flash Chat: 39.4 (#107), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Math14421414

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Expert14541421
Vectara Hallucination Rate—14.5%

Multimodal Not comparable

Longcat Flash Chat: —, Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Vision—1221

Multilingual Too close to call

Longcat Flash Chat: 51.9 (#101), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Non-English14041413
LMArena Chinese14651447
LMArena French14561455
LMArena German14081437
LMArena Japanese13731394
LMArena Korean13711384
LMArena Russian13951411
LMArena Spanish14451440

Instruction Following Too close to call

Longcat Flash Chat: 74.4 (#96), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Instruction Following14111403

Long Context Too close to call

Longcat Flash Chat: 43.5 (#93), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Longer Query14251413

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatMistral Large 3
LMArena Text14271428
LMArena Creative Writing13881386
LMArena Multi-Turn14181429
EQ-Bench Creative Writing—1412

Frequently asked questions

Is Longcat Flash Chat better than Mistral Large 3?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 39.1 on the Noometry Index.

Is Longcat Flash Chat or Mistral Large 3 better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 34.4 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Mistral Large 3 share?

19 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper