Model comparison

Gemma 3 27B vs Grok-2 (Dec 2024)

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 30.8 on the Noometry Index.

Last verified . 31 shared benchmarks.

Gemma 3 27B Google

30.8

Rank #284 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Gemma 3 27B scores higher in 4 categories and Grok-2 (Dec 2024) in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok-2 (Dec 2024) leads 38.8 to 27.6.
  • The biggest single-benchmark swing is Confabulations: 40.3% for Gemma 3 27B and 20.1% for Grok-2 (Dec 2024).
  • Gemma 3 27B has downloadable open weights; the other is API-only.

Side by side

Gemma 3 27B and Grok-2 (Dec 2024) specifications
Gemma 3 27BGrok-2 (Dec 2024)
ProviderGooglexAI
Noometry Index30.833.7
Released2025-03-112024-08-13
WeightsOpenProprietary
Context window131K—
Max output8K—
Input $ / M tokens$0.08—
Output $ / M tokens$0.16—
Results tracked4334

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok-2 (Dec 2024) leads

Gemma 3 27B: 22.5 (#334), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LiveBench Coding39.9%46.4%
LMArena Coding13221287
Aider Polyglot4.9%—
SciCode21.2%—
WeirdML—22.2%

Agentic & Tool Use Not comparable

Gemma 3 27B: 25.1 (#110), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
Berkeley Function Calling Leaderboard29.5%—

Reasoning Too close to call

Gemma 3 27B: 16.7 (#301), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LiveBench Reasoning43.8%54.8%
LMArena Hard Prompts13401272
DTBench52.5%65.2%
LiveBench Data Analysis51.5%54.5%
Epoch Capabilities Index130.04130.48
LiveBench50%54.3%
SimpleBench—22.7%
Kagi LLM Benchmark40.4%—
CritPt0%—
Chess Puzzles0%—
LMCA12.3%—

Math Gemma 3 27B leads

Gemma 3 27B: 25.9 (#265), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
OTIS Mock AIME 2024-202522.5%11.5%
LiveBench Math55.4%54.9%
LMArena Math13121283
MATH Level 574%63.5%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Grok-2 (Dec 2024) leads

Gemma 3 27B: 25.5 (#261), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
GPQA Diamond47.7%53.8%
Confabulations40.3%20.1%
LMArena Expert13041254
Vectara Hallucination Rate7.4%—

Multimodal Not comparable

Gemma 3 27B: 32.6 (#100), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LMArena Vision1164—
GeoBench52%—

Multilingual Gemma 3 27B leads

Gemma 3 27B: 46.9 (#155), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LMArena Non-English13341282
LMArena Chinese13461289
LMArena French13681318
LMArena German13621287
LMArena Japanese12871244
LMArena Korean13081237
LMArena Russian13491286
LMArena Spanish13491281

Instruction Following Gemma 3 27B leads

Gemma 3 27B: 70.6 (#160), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LiveBench Instruction Following74.9%69.6%
LMArena Instruction Following13211270

Long Context Grok-2 (Dec 2024) leads

Gemma 3 27B: 27.6 (#293), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LMArena Longer Query13331276
Fiction.LiveBench33.3%—

Writing & Preference Gemma 3 27B leads

Gemma 3 27B: 52.5 (#168), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkGemma 3 27BGrok-2 (Dec 2024)
LMArena Text13581305
LMArena Creative Writing13461284
Short-Story Creative Writing79.9%63.6%
LMArena Multi-Turn13451290
LiveBench Language34.6%45.6%
EQ-Bench Creative Writing1266—

Frequently asked questions

Is Gemma 3 27B better than Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 30.8 on the Noometry Index.

Is Gemma 3 27B or Grok-2 (Dec 2024) better for coding?

Grok-2 (Dec 2024) scores higher on coding benchmarks: 33.3 versus 22.5 in the Noometry coding category.

How many benchmarks do Gemma 3 27B and Grok-2 (Dec 2024) share?

31 benchmarks have published results for both models. Gemma 3 27B has 43 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper