Model comparison

Longcat Flash Chat vs o3

o3 is the stronger model overall, scoring 47.5 to 42.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Longcat Flash Chat scores higher in 2 categories and o3 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 40.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 43.9% for Longcat Flash Chat and 67.6% for o3.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and o3 specifications
Longcat Flash Chato3
ProviderMeituanOpenAI
Noometry Index42.147.5
Released—2025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Longcat Flash Chat: 43.5 (#87), o3: 46.8 (#64)

Coding benchmarks
BenchmarkLongcat Flash Chato3
LMArena Coding14711408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Longcat Flash Chat: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkLongcat Flash Chato3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Longcat Flash Chat: 19.0 (#272), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkLongcat Flash Chato3
Kagi LLM Benchmark43.9%67.6%
LMArena Hard Prompts14401402
ARC-AGI-2—6.5%
SimpleBench—53.1%
NYT Connections (extended)17.7%—
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Longcat Flash Chat: 39.4 (#107), o3: 50.2 (#58)

Math benchmarks
BenchmarkLongcat Flash Chato3
LMArena Math14421426
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Longcat Flash Chat: 40.6 (#116), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkLongcat Flash Chato3
LMArena Expert14541402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal Not comparable

Longcat Flash Chat: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkLongcat Flash Chato3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Too close to call

Longcat Flash Chat: 51.9 (#101), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkLongcat Flash Chato3
LMArena Non-English14041401
LMArena Chinese14651437
LMArena French14561430
LMArena German14081420
LMArena Japanese13731403
LMArena Korean13711370
LMArena Russian13951406
LMArena Spanish14451395

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkLongcat Flash Chato3
LMArena Instruction Following14111368
IFEval—86.9%

Long Context o3 leads

Longcat Flash Chat: 43.5 (#93), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkLongcat Flash Chato3
LMArena Longer Query14251372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Longcat Flash Chat: 61.0 (#91), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkLongcat Flash Chato3
LMArena Text14271410
LMArena Creative Writing13881359
LMArena Multi-Turn14181405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Longcat Flash Chat better than o3?

o3 is the stronger model overall, scoring 47.5 to 42.1 on the Noometry Index.

Is Longcat Flash Chat or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 43.5 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and o3 share?

18 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper