Model comparison

Gemma 7B vs Longcat Flash Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 30.0 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemma 7B scores higher in 1 category and Longcat Flash Chat in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 27.1.

Side by side

Gemma 7B and Longcat Flash Chat specifications
Gemma 7BLongcat Flash Chat
ProviderGoogleMeituan
Noometry Index30.042.1
Released2024-02-21—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Gemma 7B: 30.5 (#294), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Coding10481471
HumanEval+28.7%—
MBPP+43.4%—

Reasoning Too close to call

Gemma 7B: 19.9 (#249), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Hard Prompts10421440
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%
Adversarial NLI48.7%—
BIG-Bench Hard55.1%—
Epoch Capabilities Index111.99—
HellaSwag82.2%—
PIQA81.2%—
WinoGrande79%—

Math Longcat Flash Chat leads

Gemma 7B: 31.2 (#228), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Math10661442
GSM8K46.4%—

Knowledge Longcat Flash Chat leads

Gemma 7B: 27.3 (#252), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Expert10011454
ARC (AI2) Challenge78.3%—
BoolQ83.2%—
MMLU66.1%—
OpenBookQA78.6%—
TriviaQA72.3%—

Multilingual Longcat Flash Chat leads

Gemma 7B: 25.1 (#287), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Non-English9991404
LMArena Chinese10351465
LMArena French10251456
LMArena Russian9931395
LMArena German—1408
LMArena Japanese—1373
LMArena Korean—1371
LMArena Spanish—1445

Instruction Following Longcat Flash Chat leads

Gemma 7B: 51.5 (#295), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Instruction Following10171411

Long Context Longcat Flash Chat leads

Gemma 7B: 31.1 (#282), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Longer Query10221425

Writing & Preference Longcat Flash Chat leads

Gemma 7B: 27.1 (#302), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkGemma 7BLongcat Flash Chat
LMArena Text10561427
LMArena Creative Writing10241388
LMArena Multi-Turn9631418

Frequently asked questions

Is Gemma 7B better than Longcat Flash Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 30.0 on the Noometry Index.

Is Gemma 7B or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 30.5 in the Noometry coding category.

How many benchmarks do Gemma 7B and Longcat Flash Chat share?

13 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper