Model comparison

Gemma 1.1 2b IT vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 29.3 on the Noometry Index.

Last verified . 8 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 1 category and Mercury in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury leads 46.2 to 25.1.
  • Gemma 1.1 2b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 2b IT and Mercury specifications
Gemma 1.1 2b ITMercury
ProviderGoogleInception
Noometry Index29.337.6
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked169

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Gemma 1.1 2b IT: 30.1 (#299), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Coding10341322
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 19.1 (#270), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Hard Prompts10051285
Kagi LLM Benchmark—21.6%

Math Not comparable

Gemma 1.1 2b IT: 30.8 (#232), Mercury: —

Math benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Math1047—

Knowledge Not comparable

Gemma 1.1 2b IT: 26.5 (#258), Mercury: —

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Expert970—

Multilingual Mercury leads

Gemma 1.1 2b IT: 24.6 (#289), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Non-English9881260
LMArena Chinese1012—
LMArena German944—
LMArena Korean899—
LMArena Russian990—

Instruction Following Mercury leads

Gemma 1.1 2b IT: 49.9 (#299), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Instruction Following9921239

Long Context Mercury leads

Gemma 1.1 2b IT: 30.6 (#286), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Longer Query10031266

Writing & Preference Mercury leads

Gemma 1.1 2b IT: 25.1 (#306), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITMercury
LMArena Text10221282
LMArena Creative Writing9981191
LMArena Multi-Turn9591282

Frequently asked questions

Is Gemma 1.1 2b IT better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Mercury share?

8 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper