Model comparison

Gemini 3.1 Flash Lite vs GPT-5.4

GPT-5.4 is the stronger model overall, scoring 59.4 to 40.8 on the Noometry Index. Gemini 3.1 Flash Lite costs 10× less per token, which makes it the better buy when GPT-5.4's lead doesn't matter for your workload.

Last verified . 38 shared benchmarks.

Gemini 3.1 Flash Lite Google

40.8

Rank #144 Confirmed

GPT-5.4 OpenAI

59.4

Rank #16 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Gemini 3.1 Flash Lite scores higher in 0 categories and GPT-5.4 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.4 leads 61.8 to 22.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 8.2% for Gemini 3.1 Flash Lite and 91.3% for GPT-5.4.
  • Gemini 3.1 Flash Lite is cheaper at $0.25 / $1.50 per million input/output tokens, against $2.50 / $15 for GPT-5.4.
  • GPT-5.4 accepts more context: 1.05M tokens versus 1.05M.

Side by side

Gemini 3.1 Flash Lite and GPT-5.4 specifications
Gemini 3.1 Flash LiteGPT-5.4
ProviderGoogleOpenAI
Noometry Index40.859.4
Released2026-03-032026-03-05
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output66K128K
Input $ / M tokens$0.25$2.50
Output $ / M tokens$1.50$15
Results tracked3868

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 leads

Gemini 3.1 Flash Lite: 37.8 (#188), GPT-5.4: 52.6 (#33)

Coding benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
LMArena WebDev12561465
SciCode41.9%56.6%
WeirdML52.2%77.7%
LMArena Coding14001497
ALE-Bench797.731,607
SWE-bench Verified—76.9%
DeepSWE—51.8%
GSO—31.4%
MirrorCode—15.6%
AlgoTune—1.85

Agentic & Tool Use GPT-5.4 leads

Gemini 3.1 Flash Lite: 30.2 (#79), GPT-5.4: 46.5 (#13)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
DeepResearch Bench37.3%35.1%
Terminal-Bench—81.8%
APEX-Agents—52.4%
τ²-bench Banking—39.4%
PostTrainBench—19%
GBAEval—45.1%
LMArena Search—1197
METR Time Horizons—74.3%
Vending-Bench 2—6,144

Reasoning GPT-5.4 leads

Gemini 3.1 Flash Lite: 22.9 (#186), GPT-5.4: 61.8 (#19)

Reasoning benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
Kagi LLM Benchmark67.2%63.8%
NYT Connections (extended)8.2%91.3%
CritPt1.1%23.4%
Chess Puzzles25%44%
EnigmaEval3%16%
Thematic Generalization63.3%80%
LMArena Hard Prompts14071485
DTBench76.8%94.4%
LMCA35%52%
Epoch Capabilities Index144.47156.81
ForecastBench54.459.5
ARC-AGI-2—74%
ARC-AGI-1—93.7%
EBR-Bench—25.4%
Mystery Game Puzzles—37%

Math GPT-5.4 leads

Gemini 3.1 Flash Lite: 40.7 (#90), GPT-5.4: 73.5 (#19)

Knowledge GPT-5.4 leads

Gemini 3.1 Flash Lite: 41.9 (#104), GPT-5.4: 65.3 (#14)

Knowledge benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
GPQA Diamond81.8%93.3%
Humanity's Last Exam8.6%36.2%
Vectara Hallucination Rate8.2%7%
LMArena Expert13981507
SimpleQA Verified—45.1%

Multimodal GPT-5.4 leads

Gemini 3.1 Flash Lite: 39.4 (#60), GPT-5.4: 43.7 (#20)

Multimodal benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
LMArena Vision12401303
Blueprint-Bench 2—27.1%
Furniture Assembly—37.5%
LMArena Document—1471

Multilingual GPT-5.4 leads

Gemini 3.1 Flash Lite: 52.3 (#86), GPT-5.4: 56.2 (#23)

Multilingual benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
LMArena Non-English14111465
LMArena Chinese14611519
LMArena French14241493
LMArena German14291472
LMArena Japanese14131485
LMArena Korean13921448
LMArena Russian14201480
LMArena Spanish14211454

Instruction Following GPT-5.4 leads

Gemini 3.1 Flash Lite: 72.7 (#131), GPT-5.4: 77.1 (#27)

Instruction Following benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
LMArena Instruction Following13771469

Long Context GPT-5.4 leads

Gemini 3.1 Flash Lite: 42.5 (#122), GPT-5.4: 50.3 (#8)

Long Context benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
LMArena Longer Query13941473
CL-bench—27.9%
CL-bench Life—21.7%

Writing & Preference GPT-5.4 leads

Gemini 3.1 Flash Lite: 60.9 (#94), GPT-5.4: 71.9 (#17)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5.4
LMArena Text14161469
LMArena Creative Writing14011439
LMArena Multi-Turn14171482
EQ-Bench Creative Writing—1840
EQ-Bench 4—1272

Frequently asked questions

Is Gemini 3.1 Flash Lite better than GPT-5.4?

GPT-5.4 is the stronger model overall, scoring 59.4 to 40.8 on the Noometry Index. Gemini 3.1 Flash Lite costs 10× less per token, which makes it the better buy when GPT-5.4's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Flash Lite or GPT-5.4?

Gemini 3.1 Flash Lite is cheaper. It lists at $0.25 per million input tokens and $1.50 per million output tokens; GPT-5.4 lists at $2.50 and $15.

Is Gemini 3.1 Flash Lite or GPT-5.4 better for coding?

GPT-5.4 scores higher on coding benchmarks: 52.6 versus 37.8 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 does, with 1.05M tokens against 1.05M.

How many benchmarks do Gemini 3.1 Flash Lite and GPT-5.4 share?

38 benchmarks have published results for both models. Gemini 3.1 Flash Lite has 38 scored results on Noometry and GPT-5.4 has 68.

Related comparisons

Go deeper