Model comparison

Longcat Flash Chat vs Mistral Medium

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 36.3 on the Noometry Index.

Last verified . 18 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Longcat Flash Chat scores higher in 6 categories and Mistral Medium in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Longcat Flash Chat leads 40.6 to 25.0.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 43.9% for Longcat Flash Chat and 50% for Mistral Medium.

Side by side

Longcat Flash Chat and Mistral Medium specifications
Longcat Flash ChatMistral Medium
ProviderMeituanMistral AI
Noometry Index42.136.3
Released—2023-12-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Coding14711434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash ChatMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Longcat Flash Chat: 19.0 (#272), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkLongcat Flash ChatMistral Medium
Kagi LLM Benchmark43.9%50%
LMArena Hard Prompts14401426
NYT Connections (extended)17.7%—
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Math14421408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Expert14541408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Longcat Flash Chat: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Vision—1172

Multilingual Too close to call

Longcat Flash Chat: 51.9 (#101), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Non-English14041408
LMArena Chinese14651447
LMArena French14561459
LMArena German14081432
LMArena Japanese13731378
LMArena Korean13711380
LMArena Russian13951411
LMArena Spanish14451433

Instruction Following Too close to call

Longcat Flash Chat: 74.4 (#96), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Instruction Following14111398

Long Context Too close to call

Longcat Flash Chat: 43.5 (#93), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Longer Query14251406

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatMistral Medium
LMArena Text14271424
LMArena Creative Writing13881391
LMArena Multi-Turn14181418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Longcat Flash Chat better than Mistral Medium?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 36.3 on the Noometry Index.

Is Longcat Flash Chat or Mistral Medium better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 34.2 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Mistral Medium share?

18 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper