Model comparison

Grok-3 mini vs Longcat Flash Chat

Grok-3 mini and Longcat Flash Chat score almost the same on the Noometry Index (41.2 vs 42.1), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok-3 mini scores higher in 3 categories and Longcat Flash Chat in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 52.5.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 61.3% for Grok-3 mini and 43.9% for Longcat Flash Chat.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Grok-3 mini and Longcat Flash Chat specifications
Grok-3 miniLongcat Flash Chat
ProviderxAIMeituan
Noometry Index41.242.1
Released2025-04-09—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3519

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Grok-3 mini: 40.8 (#131), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Coding13791471
Aider Polyglot49.3%—
WeirdML42.6%—

Reasoning Longcat Flash Chat leads

Grok-3 mini: 13.6 (#334), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
Kagi LLM Benchmark61.3%43.9%
LMArena Hard Prompts13751440
ARC-AGI-20.4%—
NYT Connections (extended)—17.7%
ARC-AGI-116.5%—
Epoch Capabilities Index140.35—

Math Grok-3 mini leads

Grok-3 mini: 42.1 (#85), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Math13861442
OTIS Mock AIME 2024-202577.8%—
Omni-MATH31.8%—
MATH Level 590.9%—
FrontierMath (Feb 2025 set)5.9%—

Knowledge Grok-3 mini leads

Grok-3 mini: 46.4 (#81), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Expert13951454
GPQA Diamond76.3%—
MMLU-Pro79.9%—
Confabulations10.8%—
GPQA (HELM)67.5%—

Multilingual Longcat Flash Chat leads

Grok-3 mini: 48.1 (#145), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Non-English13521404
LMArena Chinese13871465
LMArena French13571456
LMArena German13491408
LMArena Japanese13421373
LMArena Korean13351371
LMArena Russian13531395
LMArena Spanish13811445

Instruction Following Grok-3 mini leads

Grok-3 mini: 78.5 (#9), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Instruction Following13571411
IFEval95.1%—

Long Context Longcat Flash Chat leads

Grok-3 mini: 41.0 (#147), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Longer Query13721425
Fiction.LiveBench66.7%—

Writing & Preference Longcat Flash Chat leads

Grok-3 mini: 52.5 (#169), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkGrok-3 miniLongcat Flash Chat
LMArena Text13701427
LMArena Creative Writing13421388
LMArena Multi-Turn13551418
Short-Story Creative Writing73.5%—
WildBench65.1%—

Frequently asked questions

Is Grok-3 mini better than Longcat Flash Chat?

Grok-3 mini and Longcat Flash Chat score almost the same on the Noometry Index (41.2 vs 42.1), so choose on price, context window or the category you care about most.

Is Grok-3 mini or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 40.8 in the Noometry coding category.

How many benchmarks do Grok-3 mini and Longcat Flash Chat share?

18 benchmarks have published results for both models. Grok-3 mini has 35 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper