Model comparison

Gemini 2.5 Flash-Lite vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 1.6× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Last verified . 22 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 2 categories and Grok 4.1 Fast in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 22.2.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 36.9% for Gemini 2.5 Flash-Lite and 69.6% for Grok 4.1 Fast.
  • Gemini 2.5 Flash-Lite is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.20 / $0.50 for Grok 4.1 Fast.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 128K.

Side by side

Gemini 2.5 Flash-Lite and Grok 4.1 Fast specifications
Gemini 2.5 Flash-LiteGrok 4.1 Fast
ProviderGooglexAI
Noometry Index37.041.4
Released2025-06-172025-06-27
WeightsProprietaryProprietary
Context window1.05M128K
Max output66K30K
Input $ / M tokens$0.10$0.20
Output $ / M tokens$0.40$0.50
Results tracked3332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.5 (#173), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Coding13731411
ALE-Bench325.9394.93
LMArena WebDev—1242
WeirdML35.2%—

Agentic & Tool Use Grok 4.1 Fast leads

Gemini 2.5 Flash-Lite: 28.0 (#96), Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
Berkeley Function Calling Leaderboard36.9%69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Gemini 2.5 Flash-Lite: 22.2 (#205), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Hard Prompts13771407
DTBench62.8%87.7%
SimpleBench—56%
Kagi LLM Benchmark40.5%—
NYT Connections (extended)—87.4%
LMCA18.1%—
Epoch Capabilities Index133.94—
ForecastBench—61

Math Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.0 (#144), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Math13731408
MathArena Final-Answer Competitions—60.9%
ProofBench—4%
Omni-MATH48%—

Knowledge Too close to call

Gemini 2.5 Flash-Lite: 32.5 (#210), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
Vectara Hallucination Rate3.3%17.8%
LMArena Expert13731399
MMLU-Pro53.7%—
GPQA (HELM)30.9%—

Multimodal Grok 4.1 Fast leads

Gemini 2.5 Flash-Lite: 29.1 (#114), Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Vision11981201
VPCT30%—

Multilingual Grok 4.1 Fast leads

Gemini 2.5 Flash-Lite: 49.3 (#134), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Non-English13691391
LMArena Chinese14041441
LMArena French13881415
LMArena German13891404
LMArena Japanese13591349
LMArena Korean13601361
LMArena Russian13731387
LMArena Spanish13961413

Instruction Following Grok 4.1 Fast leads

Gemini 2.5 Flash-Lite: 70.0 (#168), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Instruction Following13671376
IFEval81%—

Long Context Grok 4.1 Fast leads

Gemini 2.5 Flash-Lite: 33.3 (#262), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Longer Query13731390
Fiction.LiveBench47.2%—

Writing & Preference Too close to call

Gemini 2.5 Flash-Lite: 56.8 (#135), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-LiteGrok 4.1 Fast
LMArena Text13791408
LMArena Creative Writing13671394
LMArena Multi-Turn13661389
EQ-Bench Creative Writing—1327
WildBench81.8%—

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 1.6× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Which is cheaper, Gemini 2.5 Flash-Lite or Grok 4.1 Fast?

Gemini 2.5 Flash-Lite is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Is Gemini 2.5 Flash-Lite or Grok 4.1 Fast better for coding?

Gemini 2.5 Flash-Lite scores higher on coding benchmarks: 38.5 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 128K.

How many benchmarks do Gemini 2.5 Flash-Lite and Grok 4.1 Fast share?

22 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper