Model comparison

Gemini 2.5 Flash vs GPT-5.4 nano

GPT-5.4 nano is the stronger model overall, scoring 41.9 to 39.3 on the Noometry Index.

Last verified . 32 shared benchmarks.

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Gemini 2.5 Flash scores higher in 4 categories and GPT-5.4 nano in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GPT-5.4 nano leads 43.6 to 35.8.
  • The biggest single-benchmark swing is ARC-AGI-1: 33.3% for Gemini 2.5 Flash and 51.5% for GPT-5.4 nano.
  • GPT-5.4 nano is cheaper at $0.20 / $1.25 per million input/output tokens, against $0.30 / $2.50 for Gemini 2.5 Flash.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 400K.

Side by side

Gemini 2.5 Flash and GPT-5.4 nano specifications
Gemini 2.5 FlashGPT-5.4 nano
ProviderGoogleOpenAI
Noometry Index39.341.9
Released2025-04-172026-03-17
WeightsProprietaryProprietary
Context window1.05M400K
Max output66K128K
Input $ / M tokens$0.30$0.20
Output $ / M tokens$2.50$1.25
Results tracked5440

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 nano leads

Gemini 2.5 Flash: 35.8 (#220), GPT-5.4 nano: 43.6 (#84)

Coding benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
WeirdML41.9%49.2%
LMArena Coding14241405
ALE-Bench661.881,005
SWE-bench Verified (bash only)28.7%—
Aider Polyglot55.1%—
SciCode—46.9%

Agentic & Tool Use Not comparable

Gemini 2.5 Flash: 30.8 (#74), GPT-5.4 nano: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
Terminal-Bench17.1%—
Berkeley Function Calling Leaderboard56.2%—
TheAgentCompany41.1%—
BALROG33.5%—
Vending-Bench 2548.84—

Reasoning GPT-5.4 nano leads

Gemini 2.5 Flash: 18.1 (#286), GPT-5.4 nano: 23.7 (#173)

Reasoning benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
ARC-AGI-22.5%5.7%
Kagi LLM Benchmark56.8%39.7%
ARC-AGI-133.3%51.5%
CritPt1.1%9.3%
LMArena Hard Prompts14221381
DTBench76.5%80.3%
LMCA27.5%36.9%
Epoch Capabilities Index143.03145.81
ForecastBench60.657.3
SimpleBench41.2%—
Chess Puzzles—30%
EnigmaEval2.7%—
Mystery Game Puzzles—9%

Math GPT-5.4 nano leads

Gemini 2.5 Flash: 39.9 (#98), GPT-5.4 nano: 40.9 (#88)

Math benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
OTIS Mock AIME 2024-202573.1%87.8%
LMArena Math14151406
FrontierMath (Feb 2025 set)4.8%25.9%
FrontierMath Tier 4 (v1)4.2%6.3%
FrontierMath (Tiers 1-3)—44.9%
FrontierMath Tier 4—12.2%
ProofBench—5%
Omni-MATH38.5%—

Knowledge GPT-5.4 nano leads

Gemini 2.5 Flash: 36.4 (#168), GPT-5.4 nano: 41.9 (#103)

Knowledge benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
Vectara Hallucination Rate7.8%3.1%
LMArena Expert14261396
GPQA Diamond—78.5%
Humanity's Last Exam12.1%—
SimpleQA Verified—11.7%
MMLU-Pro63.9%—
Confabulations16.8%—
GPQA (HELM)39%—

Multimodal Gemini 2.5 Flash leads

Gemini 2.5 Flash: 41.8 (#32), GPT-5.4 nano: 36.7 (#78)

Multimodal benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
LMArena Vision12531196
GeoBench76%—
VPCT46.2%—
SpatialViz-Bench36.9%—

Multilingual Gemini 2.5 Flash leads

Gemini 2.5 Flash: 52.3 (#88), GPT-5.4 nano: 48.6 (#140)

Multilingual benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
LMArena Non-English14091359
LMArena Chinese14501392
LMArena French14331396
LMArena German14181367
LMArena Japanese14051343
LMArena Korean13851320
LMArena Russian14151363
LMArena Spanish14211371

Instruction Following Gemini 2.5 Flash leads

Gemini 2.5 Flash: 75.7 (#54), GPT-5.4 nano: 71.9 (#144)

Instruction Following benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
LMArena Instruction Following14051362
IFEval89.8%—

Long Context Gemini 2.5 Flash leads

Gemini 2.5 Flash: 47.5 (#17), GPT-5.4 nano: 41.6 (#137)

Long Context benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
LMArena Longer Query14191366
Fiction.LiveBench77.8%—

Writing & Preference GPT-5.4 nano leads

Gemini 2.5 Flash: 53.8 (#157), GPT-5.4 nano: 55.7 (#142)

Writing & Preference benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 nano
LMArena Text14171372
LMArena Creative Writing14001314
LMArena Multi-Turn14081382
Short-Story Creative Writing76.5%—
EQ-Bench Creative Writing1137—
WildBench81.7%—

Frequently asked questions

Is Gemini 2.5 Flash better than GPT-5.4 nano?

GPT-5.4 nano is the stronger model overall, scoring 41.9 to 39.3 on the Noometry Index.

Which is cheaper, Gemini 2.5 Flash or GPT-5.4 nano?

GPT-5.4 nano is cheaper. It lists at $0.20 per million input tokens and $1.25 per million output tokens; Gemini 2.5 Flash lists at $0.30 and $2.50.

Is Gemini 2.5 Flash or GPT-5.4 nano better for coding?

GPT-5.4 nano scores higher on coding benchmarks: 43.6 versus 35.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 400K.

How many benchmarks do Gemini 2.5 Flash and GPT-5.4 nano share?

32 benchmarks have published results for both models. Gemini 2.5 Flash has 54 scored results on Noometry and GPT-5.4 nano has 40.

Related comparisons

Go deeper