Model comparison

Gemma 4 26B A4B IT vs Grok 4.3

Gemma 4 26B A4B IT and Grok 4.3 score almost the same on the Noometry Index (43.5 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 26 shared benchmarks.

Gemma 4 26B A4B IT Google

43.5

Rank #92 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Gemma 4 26B A4B IT scores higher in 6 categories and Grok 4.3 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.3 leads 35.9 to 21.8.
  • The biggest single-benchmark swing is Chess Puzzles: 6% for Gemma 4 26B A4B IT and 25% for Grok 4.3.
  • Gemma 4 26B A4B IT is cheaper at $0.0675 / $0.23 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 262K.
  • Gemma 4 26B A4B IT has downloadable open weights; the other is API-only.

Side by side

Gemma 4 26B A4B IT and Grok 4.3 specifications
Gemma 4 26B A4B ITGrok 4.3
ProviderGooglexAI
Noometry Index43.543.8
Released2026-04-022026-04-17
WeightsOpenProprietary
Context window262K1M
Max output33K30K
Input $ / M tokens$0.0675$1.25
Output $ / M tokens$0.23$2.50
Results tracked2840

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Gemma 4 26B A4B IT: 39.0 (#164), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
LMArena WebDev13591357
SciCode40%47.3%
WeirdML35.2%49.9%
LMArena Coding14471415
ALE-Bench927.17944.17

Agentic & Tool Use Not comparable

Gemma 4 26B A4B IT: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

Gemma 4 26B A4B IT: 21.8 (#213), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
CritPt0%8%
Chess Puzzles6%25%
LMArena Hard Prompts14391396
DTBench74.9%90.7%
LMCA29.7%38.3%
Epoch Capabilities Index141.85149.16
NYT Connections (extended)—55.2%
ForecastBench—60.3

Math Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 47.6 (#67), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
OTIS Mock AIME 2024-202582.2%93.3%
LMArena Math14701388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
ProofBench—11%

Knowledge Grok 4.3 leads

Gemma 4 26B A4B IT: 45.7 (#85), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
GPQA Diamond73.2%88.8%
LMArena Expert14471385
SimpleQA Verified—33.2%
Vectara Hallucination Rate5.2%—

Multimodal Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 40.6 (#46), Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
LMArena Vision12601229
Blueprint-Bench 2—0%

Multilingual Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 53.1 (#71), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
LMArena Non-English14211385
LMArena Chinese14951422
LMArena French14601412
LMArena Russian14341399
LMArena Spanish14171398
LMArena German—1395
LMArena Japanese—1379
LMArena Korean—1356

Instruction Following Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 74.8 (#82), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
LMArena Instruction Following14201366

Long Context Gemma 4 26B A4B IT leads

Gemma 4 26B A4B IT: 43.6 (#91), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
LMArena Longer Query14281393

Writing & Preference Too close to call

Gemma 4 26B A4B IT: 58.6 (#115), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGemma 4 26B A4B ITGrok 4.3
LMArena Text14341397
LMArena Creative Writing14021380
LMArena Multi-Turn14411406
EQ-Bench Creative Writing1305—
EQ-Bench 4—1075

Frequently asked questions

Is Gemma 4 26B A4B IT better than Grok 4.3?

Gemma 4 26B A4B IT and Grok 4.3 score almost the same on the Noometry Index (43.5 vs 43.8), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 4 26B A4B IT or Grok 4.3?

Gemma 4 26B A4B IT is cheaper. It lists at $0.0675 per million input tokens and $0.23 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Gemma 4 26B A4B IT or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 39.0 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 262K.

How many benchmarks do Gemma 4 26B A4B IT and Grok 4.3 share?

26 benchmarks have published results for both models. Gemma 4 26B A4B IT has 28 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper