Model comparison

Gemma 2 9B vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 25.9 on the Noometry Index.

Last verified . 20 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemma 2 9B scores higher in 0 categories and Grok 4.3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 9.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 0.6% for Gemma 2 9B and 93.3% for Grok 4.3.
  • Gemma 2 9B has downloadable open weights; the other is API-only.

Side by side

Gemma 2 9B and Grok 4.3 specifications
Gemma 2 9BGrok 4.3
ProviderGooglexAI
Noometry Index25.943.8
Released2024-06-242026-04-17
WeightsOpenProprietary
Context window—1M
Max output—30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked3540

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Gemma 2 9B: 29.4 (#304), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Coding11731415
LMArena WebDev—1357
SciCode—47.3%
WeirdML—49.9%
BigCodeBench Instruct34.7%—
LiveBench Coding22.5%—
BigCodeBench Complete40.6%—
ALE-Bench—944.17

Agentic & Tool Use Not comparable

Gemma 2 9B: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 9BGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

Gemma 2 9B: 15.9 (#309), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Hard Prompts11711396
Epoch Capabilities Index119.83149.16
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%
LiveBench Reasoning15.2%—
DTBench—90.7%
LiveBench Data Analysis36.4%—
LMCA—38.3%
ForecastBench—60.3
LiveBench28.7%—
PIQA83.7%—

Math Grok 4.3 leads

Gemma 2 9B: 9.9 (#318), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkGemma 2 9BGrok 4.3
OTIS Mock AIME 2024-20250.6%93.3%
LMArena Math11831388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
ProofBench—11%
LiveBench Math19.8%—
MATH Level 521%—
GSM8K84.9%—

Knowledge Grok 4.3 leads

Gemma 2 9B: 9.7 (#305), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGemma 2 9BGrok 4.3
GPQA Diamond27.5%88.8%
LMArena Expert11471385
SimpleQA Verified—33.2%
BoolQ85.7%—
MMLU72.1%—

Multimodal Not comparable

Gemma 2 9B: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual Grok 4.3 leads

Gemma 2 9B: 36.6 (#238), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Non-English11881385
LMArena Chinese11851422
LMArena French11901412
LMArena German11861395
LMArena Japanese11441379
LMArena Korean11371356
LMArena Russian12001399
LMArena Spanish12001398

Instruction Following Grok 4.3 leads

Gemma 2 9B: 57.6 (#269), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Instruction Following11781366
LiveBench Instruction Following52.6%—

Long Context Grok 4.3 leads

Gemma 2 9B: 36.3 (#233), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Longer Query11971393

Writing & Preference Grok 4.3 leads

Gemma 2 9B: 32.1 (#281), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGemma 2 9BGrok 4.3
LMArena Text12071397
LMArena Creative Writing12061380
LMArena Multi-Turn11931406
EQ-Bench Creative Writing841—
EQ-Bench 4—1075
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 25.9 on the Noometry Index.

Is Gemma 2 9B or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and Grok 4.3 share?

20 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper