Model comparison

Gemma 2B vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 29.6 on the Noometry Index.

Last verified . 11 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Gemma 2B scores higher in 0 categories and Grok 4.1 Fast in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 Fast leads 57.2 to 24.0.
  • Gemma 2B has downloadable open weights; the other is API-only.

Side by side

Gemma 2B and Grok 4.1 Fast specifications
Gemma 2BGrok 4.1 Fast
ProviderGooglexAI
Noometry Index29.641.4
Released2024-02-212025-06-27
WeightsOpenProprietary
Context window—128K
Max output—30K
Input $ / M tokens—$0.20
Output $ / M tokens—$0.50
Results tracked2332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 Fast leads

Gemma 2B: 29.4 (#305), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Coding10101411
LMArena WebDev—1242
ALE-Bench—394.93
HumanEval+20.7%—
MBPP+34.1%—

Agentic & Tool Use Not comparable

Gemma 2B: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Gemma 2B: 18.8 (#275), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Hard Prompts9891407
SimpleBench—56%
NYT Connections (extended)—87.4%
DTBench—87.7%
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
ForecastBench—61
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Grok 4.1 Fast leads

Gemma 2B: 30.0 (#239), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Math10091408
MathArena Final-Answer Competitions—60.9%
ProofBench—4%
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
Vectara Hallucination Rate—17.8%
LMArena Expert—1399
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multimodal Not comparable

Gemma 2B: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Vision—1201

Multilingual Grok 4.1 Fast leads

Gemma 2B: 23.0 (#294), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Non-English9581391
LMArena Chinese9861441
LMArena Russian9371387
LMArena French—1415
LMArena German—1404
LMArena Japanese—1349
LMArena Korean—1361
LMArena Spanish—1413

Instruction Following Grok 4.1 Fast leads

Gemma 2B: 48.5 (#302), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Instruction Following9701376

Long Context Grok 4.1 Fast leads

Gemma 2B: 29.9 (#291), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Longer Query9811390

Writing & Preference Grok 4.1 Fast leads

Gemma 2B: 24.0 (#308), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGemma 2BGrok 4.1 Fast
LMArena Text10021408
LMArena Creative Writing9871394
LMArena Multi-Turn9451389
EQ-Bench Creative Writing—1327

Frequently asked questions

Is Gemma 2B better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 29.6 on the Noometry Index.

Is Gemma 2B or Grok 4.1 Fast better for coding?

Grok 4.1 Fast scores higher on coding benchmarks: 34.1 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and Grok 4.1 Fast share?

11 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper