Model comparison

Gemma 1.1 2b IT vs Grok Build 0.1

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 29.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 19.1.
  • Gemma 1.1 2b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 2b IT and Grok Build 0.1 specifications
Gemma 1.1 2b ITGrok Build 0.1
ProviderGooglexAI
Noometry Index29.336.4
Released—2026-04-16
WeightsOpenProprietary
Context window—256K
Max output—256K
Input $ / M tokens—$1
Output $ / M tokens—$2
Results tracked163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Gemma 1.1 2b IT: 30.1 (#299), Grok Build 0.1: 43.1 (#91)

Coding benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
SciCode—50.2%
LMArena Coding1034—
HumanEval+17.7%—
MBPP+23.3%—

Agentic & Tool Use Not comparable

Gemma 1.1 2b IT: —, Grok Build 0.1: 22.7 (#129)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
GBAEval—2.4%

Reasoning Grok Build 0.1 leads

Gemma 1.1 2b IT: 19.1 (#270), Grok Build 0.1: 32.2 (#77)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
CritPt—9.1%
LMArena Hard Prompts1005—

Math Not comparable

Gemma 1.1 2b IT: 30.8 (#232), Grok Build 0.1: —

Math benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
LMArena Math1047—

Knowledge Not comparable

Gemma 1.1 2b IT: 26.5 (#258), Grok Build 0.1: —

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
LMArena Expert970—

Multilingual Not comparable

Gemma 1.1 2b IT: 24.6 (#289), Grok Build 0.1: —

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
LMArena Non-English988—
LMArena Chinese1012—
LMArena German944—
LMArena Korean899—
LMArena Russian990—

Instruction Following Not comparable

Gemma 1.1 2b IT: 49.9 (#299), Grok Build 0.1: —

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
LMArena Instruction Following992—

Long Context Not comparable

Gemma 1.1 2b IT: 30.6 (#286), Grok Build 0.1: —

Long Context benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
LMArena Longer Query1003—

Writing & Preference Not comparable

Gemma 1.1 2b IT: 25.1 (#306), Grok Build 0.1: —

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITGrok Build 0.1
LMArena Text1022—
LMArena Creative Writing998—
LMArena Multi-Turn959—

Frequently asked questions

Is Gemma 1.1 2b IT better than Grok Build 0.1?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Grok Build 0.1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Grok Build 0.1 share?

0 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Grok Build 0.1 has 3.

Related comparisons

Go deeper