Model comparison

Longcat Flash Chat vs Mistral Large

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Longcat Flash Chat scores higher in 8 categories and Mistral Large in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Longcat Flash Chat leads 39.4 to 18.2.

Side by side

Longcat Flash Chat and Mistral Large specifications
Longcat Flash ChatMistral Large
ProviderMeituanMistral AI
Noometry Index42.131.9
Released—2024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Coding14711277
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Longcat Flash Chat leads

Longcat Flash Chat: 19.0 (#272), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Hard Prompts14401257
SimpleBench—22.5%
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Math14421262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Expert14541232
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Non-English14041237
LMArena Chinese14651240
LMArena French14561325
LMArena German14081254
LMArena Japanese13731188
LMArena Korean13711202
LMArena Russian13951257
LMArena Spanish14451268

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Instruction Following14111249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Longer Query14251261

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatMistral Large
LMArena Text14271266
LMArena Creative Writing13881243
LMArena Multi-Turn14181260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Longcat Flash Chat better than Mistral Large?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 31.9 on the Noometry Index.

Is Longcat Flash Chat or Mistral Large better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 34.3 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Mistral Large share?

17 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper