Model comparison

Gemini 3.5 Flash Lite vs Grok 4

Grok 4 is the stronger model overall, scoring 48.1 to 41.5 on the Noometry Index.

Last verified . 25 shared benchmarks.

Gemini 3.5 Flash Lite Google

41.5

Rank #133 Confirmed

Grok 4 xAI

48.1

Rank #56 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Gemini 3.5 Flash Lite scores higher in 3 categories and Grok 4 in 7 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4 leads 48.4 to 25.9.
  • The biggest single-benchmark swing is ARC-AGI-1: 53.5% for Gemini 3.5 Flash Lite and 66.7% for Grok 4.

Side by side

Gemini 3.5 Flash Lite and Grok 4 specifications
Gemini 3.5 Flash LiteGrok 4
ProviderGooglexAI
Noometry Index41.548.1
Released2026-07-212025-07-09
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$0.30—
Output $ / M tokens$2.50—
Results tracked3948

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4 leads

Gemini 3.5 Flash Lite: 42.4 (#104), Grok 4: 50.3 (#46)

Coding benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
WeirdML39%45.7%
LMArena Coding14531408
Aider Polyglot—79.6%
LMArena WebDev1440—
SciCode41.3%—
ALE-Bench765.27—

Agentic & Tool Use Grok 4 leads

Gemini 3.5 Flash Lite: 25.7 (#107), Grok 4: 32.3 (#68)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
Terminal-Bench—27.2%
APEX-Agents29.3%—
Berkeley Function Calling Leaderboard—63%
GDPval—21.1%
Cybench—43%
DeepResearch Bench—47.3%
BALROG—43.6%
GDP.pdf10%—
LMArena Search—1142
METR Time Horizons—66.6%

Reasoning Grok 4 leads

Gemini 3.5 Flash Lite: 27.8 (#114), Grok 4: 36.7 (#65)

Reasoning benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
ARC-AGI-210.3%16%
ARC-AGI-153.5%66.7%
Chess Puzzles22%28%
LMArena Hard Prompts14411409
Epoch Capabilities Index145.13146.44
SimpleBench—60.5%
Kagi LLM Benchmark—73.6%
NYT Connections (extended)60.4%—
CritPt0%—
Mystery Game Puzzles19%—
DTBench83.5%—
LMCA37.6%—
ForecastBench—60.9

Math Grok 4 leads

Gemini 3.5 Flash Lite: 25.9 (#264), Grok 4: 48.4 (#64)

Math benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
OTIS Mock AIME 2024-202571.1%84%
LMArena Math14331422
FrontierMath (Tiers 1-3)26%—
FrontierMath Tier 40%—
ProofBench13%—
Omni-MATH—60.3%
FrontierMath (Feb 2025 set)—19.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Grok 4 leads

Gemini 3.5 Flash Lite: 50.1 (#72), Grok 4: 53.8 (#55)

Knowledge benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
GPQA Diamond83.3%87%
LMArena Expert14351415
MMLU-Pro—85.1%
Confabulations—12.4%
GPQA (HELM)—72.7%

Multimodal Gemini 3.5 Flash Lite leads

Gemini 3.5 Flash Lite: 41.1 (#40), Grok 4: 33.7 (#94)

Multimodal benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
LMArena Vision12681210
GeoBench—45%

Multilingual Gemini 3.5 Flash Lite leads

Gemini 3.5 Flash Lite: 53.5 (#63), Grok 4: 51.8 (#103)

Multilingual benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
LMArena Non-English14271403
LMArena Chinese14691427
LMArena French14511418
LMArena German14541429
LMArena Japanese14271394
LMArena Korean14071377
LMArena Russian14431410
LMArena Spanish14421420

Instruction Following Grok 4 leads

Gemini 3.5 Flash Lite: 74.8 (#85), Grok 4: 79.2 (#5)

Instruction Following benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
LMArena Instruction Following14201387
IFEval—94.9%

Long Context Grok 4 leads

Gemini 3.5 Flash Lite: 43.9 (#84), Grok 4: 63.1 (#4)

Long Context benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
LMArena Longer Query14351409
Fiction.LiveBench—94.4%

Writing & Preference Gemini 3.5 Flash Lite leads

Gemini 3.5 Flash Lite: 64.2 (#57), Grok 4: 58.5 (#116)

Writing & Preference benchmarks
BenchmarkGemini 3.5 Flash LiteGrok 4
LMArena Text14351411
LMArena Creative Writing14201397
LMArena Multi-Turn14451416
Short-Story Creative Writing—76.9%
EQ-Bench Creative Writing1559—
WildBench—79.7%

Frequently asked questions

Is Gemini 3.5 Flash Lite better than Grok 4?

Grok 4 is the stronger model overall, scoring 48.1 to 41.5 on the Noometry Index.

Is Gemini 3.5 Flash Lite or Grok 4 better for coding?

Grok 4 scores higher on coding benchmarks: 50.3 versus 42.4 in the Noometry coding category.

How many benchmarks do Gemini 3.5 Flash Lite and Grok 4 share?

25 benchmarks have published results for both models. Gemini 3.5 Flash Lite has 39 scored results on Noometry and Grok 4 has 48.

Related comparisons

Go deeper