Model comparison

Longcat Flash Chat vs Phi 3 Mini 4k Instruct June 2024

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 31.3 on the Noometry Index.

Last verified . 15 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Longcat Flash Chat scores higher in 7 categories and Phi 3 Mini 4k Instruct June 2024 in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 29.6.

Side by side

Longcat Flash Chat and Phi 3 Mini 4k Instruct June 2024 specifications
Longcat Flash ChatPhi 3 Mini 4k Instruct June 2024
ProviderMeituanMicrosoft
Noometry Index42.131.3
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#87), Phi 3 Mini 4k Instruct June 2024: 31.8 (#279)

Coding benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Coding14711093

Reasoning Phi 3 Mini 4k Instruct June 2024 leads

Longcat Flash Chat: 19.0 (#272), Phi 3 Mini 4k Instruct June 2024: 20.8 (#231)

Reasoning benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Hard Prompts14401087
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—

Math Longcat Flash Chat leads

Longcat Flash Chat: 39.4 (#107), Phi 3 Mini 4k Instruct June 2024: 33.0 (#208)

Math benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Math14421152

Knowledge Longcat Flash Chat leads

Longcat Flash Chat: 40.6 (#116), Phi 3 Mini 4k Instruct June 2024: 28.6 (#244)

Knowledge benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Expert14541051

Multilingual Longcat Flash Chat leads

Longcat Flash Chat: 51.9 (#101), Phi 3 Mini 4k Instruct June 2024: 25.9 (#282)

Multilingual benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Non-English14041013
LMArena Chinese14651033
LMArena German14081031
LMArena Japanese1373954
LMArena Korean1371880
LMArena Russian13951019
LMArena French1456—
LMArena Spanish1445—

Instruction Following Longcat Flash Chat leads

Longcat Flash Chat: 74.4 (#96), Phi 3 Mini 4k Instruct June 2024: 54.1 (#282)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Instruction Following14111058

Long Context Longcat Flash Chat leads

Longcat Flash Chat: 43.5 (#93), Phi 3 Mini 4k Instruct June 2024: 31.7 (#277)

Long Context benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Longer Query14251042

Writing & Preference Longcat Flash Chat leads

Longcat Flash Chat: 61.0 (#91), Phi 3 Mini 4k Instruct June 2024: 29.6 (#294)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatPhi 3 Mini 4k Instruct June 2024
LMArena Text14271080
LMArena Creative Writing13881045
LMArena Multi-Turn14181049

Frequently asked questions

Is Longcat Flash Chat better than Phi 3 Mini 4k Instruct June 2024?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 31.3 on the Noometry Index.

Is Longcat Flash Chat or Phi 3 Mini 4k Instruct June 2024 better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 31.8 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Phi 3 Mini 4k Instruct June 2024 share?

15 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Phi 3 Mini 4k Instruct June 2024 has 15.

Related comparisons

Go deeper