Model comparison

Gemma 3 4B vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 28.1 on the Noometry Index. Gemma 3 4B costs 5.5× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Last verified . 16 shared benchmarks.

Gemma 3 4B Google

28.1

Rank #326 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Gemma 3 4B scores higher in 1 category and Grok 4.1 Fast in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 13.2.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 19.6% for Gemma 3 4B and 69.6% for Grok 4.1 Fast.
  • Gemma 3 4B is cheaper at $0.04 / $0.08 per million input/output tokens, against $0.20 / $0.50 for Grok 4.1 Fast.
  • Gemma 3 4B accepts more context: 131K tokens versus 128K.
  • Gemma 3 4B has downloadable open weights; the other is API-only.

Side by side

Gemma 3 4B and Grok 4.1 Fast specifications
Gemma 3 4BGrok 4.1 Fast
ProviderGooglexAI
Noometry Index28.141.4
Released2025-03-122025-06-27
WeightsOpenProprietary
Context window131K128K
Max output4K30K
Input $ / M tokens$0.04$0.20
Output $ / M tokens$0.08$0.50
Results tracked2232

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3 4B leads

Gemma 3 4B: 35.9 (#215), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Coding12301411
LMArena WebDev—1242
ALE-Bench—394.93

Agentic & Tool Use Grok 4.1 Fast leads

Gemma 3 4B: 20.9 (#142), Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
Berkeley Function Calling Leaderboard19.6%69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Gemma 3 4B: 13.2 (#335), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Hard Prompts12531407
DTBench50.9%87.7%
SimpleBench—56%
Kagi LLM Benchmark25.2%—
NYT Connections (extended)—87.4%
Chess Puzzles0%—
LMCA2.8%—
Epoch Capabilities Index116.02—
ForecastBench—61

Math Grok 4.1 Fast leads

Gemma 3 4B: 16.8 (#292), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Math12391408
MathArena Final-Answer Competitions—60.9%
OTIS Mock AIME 2024-20257.5%—
ProofBench—4%

Knowledge Grok 4.1 Fast leads

Gemma 3 4B: 11.8 (#299), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
Vectara Hallucination Rate6.4%17.8%
LMArena Expert12231399
GPQA Diamond23.2%—

Multimodal Not comparable

Gemma 3 4B: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Vision—1201

Multilingual Grok 4.1 Fast leads

Gemma 3 4B: 42.5 (#194), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Non-English12731391
LMArena German12811404
LMArena Russian12941387
LMArena Chinese—1441
LMArena French—1415
LMArena Japanese—1349
LMArena Korean—1361
LMArena Spanish—1413

Instruction Following Grok 4.1 Fast leads

Gemma 3 4B: 65.2 (#225), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Instruction Following12391376

Long Context Grok 4.1 Fast leads

Gemma 3 4B: 38.7 (#194), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Longer Query12731390

Writing & Preference Grok 4.1 Fast leads

Gemma 3 4B: 42.0 (#239), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGemma 3 4BGrok 4.1 Fast
LMArena Text12911408
LMArena Creative Writing12711394
EQ-Bench Creative Writing10681327
LMArena Multi-Turn12551389

Frequently asked questions

Is Gemma 3 4B better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 28.1 on the Noometry Index. Gemma 3 4B costs 5.5× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Which is cheaper, Gemma 3 4B or Grok 4.1 Fast?

Gemma 3 4B is cheaper. It lists at $0.04 per million input tokens and $0.08 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Is Gemma 3 4B or Grok 4.1 Fast better for coding?

Gemma 3 4B scores higher on coding benchmarks: 35.9 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Gemma 3 4B does, with 131K tokens against 128K.

How many benchmarks do Gemma 3 4B and Grok 4.1 Fast share?

16 benchmarks have published results for both models. Gemma 3 4B has 22 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper