Model comparison

Claude 3.5 Sonnet vs Longcat Flash Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 34.6 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 1 category and Longcat Flash Chat in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Longcat Flash Chat leads 39.4 to 19.2.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Longcat Flash Chat specifications
Claude 3.5 SonnetLongcat Flash Chat
ProviderAnthropicMeituan
Noometry Index34.642.1
Released2024-06-20—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Claude 3.5 Sonnet: 39.0 (#165), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Coding13421471
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Longcat Flash Chat: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Hard Prompts13051440
SimpleBench41.4%—
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Longcat Flash Chat leads

Claude 3.5 Sonnet: 19.2 (#288), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Math13071442
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Longcat Flash Chat leads

Claude 3.5 Sonnet: 28.6 (#245), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Expert12651454
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Longcat Flash Chat: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Longcat Flash Chat leads

Claude 3.5 Sonnet: 43.2 (#185), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Non-English12831404
LMArena Chinese12721465
LMArena French13051456
LMArena German12971408
LMArena Japanese12341373
LMArena Korean12001371
LMArena Russian13061395
LMArena Spanish12901445

Instruction Following Longcat Flash Chat leads

Claude 3.5 Sonnet: 68.8 (#182), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Instruction Following12971411
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Longcat Flash Chat leads

Claude 3.5 Sonnet: 39.9 (#167), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Longer Query13111425

Writing & Preference Longcat Flash Chat leads

Claude 3.5 Sonnet: 52.9 (#164), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetLongcat Flash Chat
LMArena Text12981427
LMArena Creative Writing12921388
LMArena Multi-Turn13261418
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Longcat Flash Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Longcat Flash Chat share?

17 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper