Model comparison

Gemma 1.1 2b IT vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 29.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 0 categories and Muse Spark in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Muse Spark leads 66.0 to 25.1.
  • Gemma 1.1 2b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 2b IT and Muse Spark specifications
Gemma 1.1 2b ITMuse Spark
ProviderGoogleMeta
Noometry Index29.350.6
Released—2026-04-08
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1627

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Gemma 1.1 2b IT: 30.1 (#299), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Coding10341481
SciCode—51.5%
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Muse Spark leads

Gemma 1.1 2b IT: 19.1 (#270), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Hard Prompts10051474
CritPt—11.3%
Epoch Capabilities Index—152.04

Math Muse Spark leads

Gemma 1.1 2b IT: 30.8 (#232), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Math10471455
OTIS Mock AIME 2024-2025—88.9%
ProofBench—17%
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Gemma 1.1 2b IT: 26.5 (#258), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Expert9701457
GPQA Diamond—89.8%
Humanity's Last Exam—40.6%

Multimodal Not comparable

Gemma 1.1 2b IT: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Muse Spark leads

Gemma 1.1 2b IT: 24.6 (#289), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Non-English9881464
LMArena Chinese10121509
LMArena German9441497
LMArena Korean8991459
LMArena Russian9901466
LMArena French—1497
LMArena Spanish—1472

Instruction Following Muse Spark leads

Gemma 1.1 2b IT: 49.9 (#299), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Instruction Following9921442

Long Context Muse Spark leads

Gemma 1.1 2b IT: 30.6 (#286), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Longer Query10031451

Writing & Preference Muse Spark leads

Gemma 1.1 2b IT: 25.1 (#306), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITMuse Spark
LMArena Text10221474
LMArena Creative Writing9981459
LMArena Multi-Turn9591477

Frequently asked questions

Is Gemma 1.1 2b IT better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Muse Spark share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper