Model comparison

Longcat Flash Chat vs o3-pro

Longcat Flash Chat and o3-pro score almost the same on the Noometry Index (42.1 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • They share 1 benchmark with published results for both. Longcat Flash Chat scores higher in 2 categories and o3-pro in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3-pro leads 72.2 to 43.5.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 43.9% for Longcat Flash Chat and 72.1% for o3-pro.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and o3-pro specifications
Longcat Flash Chato3-pro
ProviderMeituanOpenAI
Noometry Index42.142.9
Released—2025-06-10
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$20
Output $ / M tokens—$80
Results tracked1912

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-pro leads

Longcat Flash Chat: 43.5 (#87), o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkLongcat Flash Chato3-pro
Aider Polyglot—84.9%
WeirdML—58.2%
LMArena Coding1471—

Reasoning o3-pro leads

Longcat Flash Chat: 19.0 (#272), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkLongcat Flash Chato3-pro
Kagi LLM Benchmark43.9%72.1%
ARC-AGI-2—4.9%
NYT Connections (extended)17.7%—
ARC-AGI-1—59.3%
LMArena Hard Prompts1440—
DTBench—86.9%
LMCA—38.5%
Epoch Capabilities Index—147.42

Math Not comparable

Longcat Flash Chat: 39.4 (#107), o3-pro: —

Math benchmarks
BenchmarkLongcat Flash Chato3-pro
LMArena Math1442—

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkLongcat Flash Chato3-pro
Confabulations—14.2%
Vectara Hallucination Rate—23.3%
LMArena Expert1454—

Multilingual Not comparable

Longcat Flash Chat: 51.9 (#101), o3-pro: —

Multilingual benchmarks
BenchmarkLongcat Flash Chato3-pro
LMArena Non-English1404—
LMArena Chinese1465—
LMArena French1456—
LMArena German1408—
LMArena Japanese1373—
LMArena Korean1371—
LMArena Russian1395—
LMArena Spanish1445—

Instruction Following Not comparable

Longcat Flash Chat: 74.4 (#96), o3-pro: —

Instruction Following benchmarks
BenchmarkLongcat Flash Chato3-pro
LMArena Instruction Following1411—

Long Context o3-pro leads

Longcat Flash Chat: 43.5 (#93), o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkLongcat Flash Chato3-pro
Fiction.LiveBench—97.2%
LMArena Longer Query1425—

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkLongcat Flash Chato3-pro
LMArena Text1427—
LMArena Creative Writing1388—
Short-Story Creative Writing—84.4%
LMArena Multi-Turn1418—

Frequently asked questions

Is Longcat Flash Chat better than o3-pro?

Longcat Flash Chat and o3-pro score almost the same on the Noometry Index (42.1 vs 42.9), so choose on price, context window or the category you care about most.

Is Longcat Flash Chat or o3-pro better for coding?

o3-pro scores higher on coding benchmarks: 55.5 versus 43.5 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and o3-pro share?

1 benchmark has published results for both models. Longcat Flash Chat has 19 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper