Model comparison

Gemini 2.5 Flash-Lite vs GPT-4o

Gemini 2.5 Flash-Lite is the stronger model overall, scoring 37.0 to 28.6 on the Noometry Index.

Last verified . 30 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

GPT-4o OpenAI

28.6

Rank #324 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 8 categories and GPT-4o in 2 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 2.5 Flash-Lite leads 38.0 to 10.6.
  • The biggest single-benchmark swing is GPQA (HELM): 30.9% for Gemini 2.5 Flash-Lite and 52% for GPT-4o.
  • Gemini 2.5 Flash-Lite is cheaper at $0.10 / $0.40 per million input/output tokens, against $2.50 / $10 for GPT-4o.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 128K.

Side by side

Gemini 2.5 Flash-Lite and GPT-4o specifications
Gemini 2.5 Flash-LiteGPT-4o
ProviderGoogleOpenAI
Noometry Index37.028.6
Released2025-06-172024-05-13
WeightsProprietaryProprietary
Context window1.05M128K
Max output66K16K
Input $ / M tokens$0.10$2.50
Output $ / M tokens$0.40$10
Results tracked3372

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.5 (#173), GPT-4o: 24.8 (#328)

Coding benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
WeirdML35.2%25.1%
LMArena Coding13731297
SWE-bench Verified—31%
SWE-bench Verified (bash only)—21.6%
Aider Polyglot—45.3%
GSO—0%
BigCodeBench Instruct—51.1%
LiveBench Coding—51.4%
BigCodeBench Complete—61.1%
CadEval—26%
ALE-Bench325.9—
HumanEval+—87.2%
MBPP+—72.2%

Agentic & Tool Use Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 28.0 (#96), GPT-4o: 21.0 (#141)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
Berkeley Function Calling Leaderboard36.9%—
GDPval—9.9%
TheAgentCompany—8.6%
Cybench—12.5%
BALROG—32.3%
LMArena Search—1006
METR Time Horizons—40.8%

Reasoning Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 22.2 (#205), GPT-4o: 9.4 (#343)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
LMArena Hard Prompts13771281
DTBench62.8%64.5%
LMCA18.1%16.6%
Epoch Capabilities Index133.94128.97
ARC-AGI-2—0%
SimpleBench—17.8%
Kagi LLM Benchmark40.5%—
ARC-AGI-1—4.5%
CritPt—0%
Chess Puzzles—13%
EnigmaEval—0.8%
LiveBench Reasoning—55.8%
LiveBench Data Analysis—60.9%
ForecastBench—57.7
LiveBench—55.3%

Math Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.0 (#144), GPT-4o: 10.6 (#312)

Math benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
Omni-MATH48%29.3%
LMArena Math13731285
FrontierMath (Tiers 1-3)—0.4%
OTIS Mock AIME 2024-2025—6.4%
LiveBench Math—49.5%
MATH Level 5—53.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 32.5 (#210), GPT-4o: 28.8 (#242)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
MMLU-Pro53.7%71.3%
Vectara Hallucination Rate3.3%9.6%
GPQA (HELM)30.9%52%
LMArena Expert13731250
GPQA Diamond—49.2%
Humanity's Last Exam—2.7%
SimpleQA Verified—26%
Confabulations—15.3%
MMLU—88.1%

Multimodal GPT-4o leads

Gemini 2.5 Flash-Lite: 29.1 (#114), GPT-4o: 34.5 (#91)

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
LMArena Vision11981137
VPCT30%40%
Video-MME—71.9%
GeoBench—71%
ScienceQA—88.5%

Multilingual Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 49.3 (#134), GPT-4o: 43.2 (#186)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
LMArena Non-English13691283
LMArena Chinese14041277
LMArena French13881304
LMArena German13891282
LMArena Japanese13591257
LMArena Korean13601234
LMArena Russian13731286
LMArena Spanish13961292

Instruction Following Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 70.0 (#168), GPT-4o: 66.6 (#207)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
IFEval81%81.7%
LMArena Instruction Following13671278
LiveBench Instruction Following—68.6%

Long Context GPT-4o leads

Gemini 2.5 Flash-Lite: 33.3 (#262), GPT-4o: 39.4 (#179)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
Fiction.LiveBench47.2%66.7%
LMArena Longer Query13731289

Writing & Preference Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 56.8 (#135), GPT-4o: 52.6 (#166)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-LiteGPT-4o
LMArena Text13791300
LMArena Creative Writing13671292
WildBench81.8%82.8%
LMArena Multi-Turn13661302
Short-Story Creative Writing—81.8%
LiveBench Language—47.6%

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than GPT-4o?

Gemini 2.5 Flash-Lite is the stronger model overall, scoring 37.0 to 28.6 on the Noometry Index.

Which is cheaper, Gemini 2.5 Flash-Lite or GPT-4o?

Gemini 2.5 Flash-Lite is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; GPT-4o lists at $2.50 and $10.

Is Gemini 2.5 Flash-Lite or GPT-4o better for coding?

Gemini 2.5 Flash-Lite scores higher on coding benchmarks: 38.5 versus 24.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 128K.

How many benchmarks do Gemini 2.5 Flash-Lite and GPT-4o share?

30 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and GPT-4o has 72.

Related comparisons

Go deeper