Model comparison

Inkling vs Longcat Flash Chat

Inkling is the stronger model overall, scoring 44.1 to 42.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Inkling scores higher in 6 categories and Longcat Flash Chat in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 19.0.

Side by side

Inkling and Longcat Flash Chat specifications
InklingLongcat Flash Chat
ProviderThinking Machines LabMeituan
Noometry Index44.142.1
Released2026-07-15—
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4119

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Inkling: 34.5 (#234), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Coding14641471
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Longcat Flash Chat: —

Agentic & Tool Use benchmarks
BenchmarkInklingLongcat Flash Chat
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Hard Prompts14511440
ARC-AGI-236.5%—
SimpleBench50%—
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Longcat Flash Chat leads

Inkling: 31.3 (#225), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Math14791442
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Expert14651454
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Non-English14341404
LMArena Chinese14901465
LMArena French14581456
LMArena German14461408
LMArena Japanese14291373
LMArena Korean14041371
LMArena Russian14291395
LMArena Spanish14481445

Instruction Following Too close to call

Inkling: 75.1 (#71), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Instruction Following14261411

Long Context Too close to call

Inkling: 43.8 (#86), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Longer Query14341425

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkInklingLongcat Flash Chat
LMArena Text14411427
LMArena Creative Writing13871388
LMArena Multi-Turn14361418
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Longcat Flash Chat?

Inkling is the stronger model overall, scoring 44.1 to 42.1 on the Noometry Index.

Is Inkling or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Longcat Flash Chat share?

17 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper