Model comparison

Gemma 1.1 2b IT vs Longcat Flash Chat

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 29.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 1 category and Longcat Flash Chat in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Longcat Flash Chat leads 61.0 to 25.1.

Side by side

Gemma 1.1 2b IT and Longcat Flash Chat specifications
Gemma 1.1 2b ITLongcat Flash Chat
ProviderGoogleMeituan
Noometry Index29.342.1
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Longcat Flash Chat leads

Gemma 1.1 2b IT: 30.1 (#299), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Coding10341471
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Too close to call

Gemma 1.1 2b IT: 19.1 (#270), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Hard Prompts10051440
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%

Math Longcat Flash Chat leads

Gemma 1.1 2b IT: 30.8 (#232), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Math10471442

Knowledge Longcat Flash Chat leads

Gemma 1.1 2b IT: 26.5 (#258), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Expert9701454

Multilingual Longcat Flash Chat leads

Gemma 1.1 2b IT: 24.6 (#289), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Non-English9881404
LMArena Chinese10121465
LMArena German9441408
LMArena Korean8991371
LMArena Russian9901395
LMArena French—1456
LMArena Japanese—1373
LMArena Spanish—1445

Instruction Following Longcat Flash Chat leads

Gemma 1.1 2b IT: 49.9 (#299), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Instruction Following9921411

Long Context Longcat Flash Chat leads

Gemma 1.1 2b IT: 30.6 (#286), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Longer Query10031425

Writing & Preference Longcat Flash Chat leads

Gemma 1.1 2b IT: 25.1 (#306), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITLongcat Flash Chat
LMArena Text10221427
LMArena Creative Writing9981388
LMArena Multi-Turn9591418

Frequently asked questions

Is Gemma 1.1 2b IT better than Longcat Flash Chat?

Longcat Flash Chat is the stronger model overall, scoring 42.1 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Longcat Flash Chat better for coding?

Longcat Flash Chat scores higher on coding benchmarks: 43.5 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Longcat Flash Chat share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper