Model comparison

Gemini 2.5 Flash vs GPT-5.4 mini

GPT-5.4 mini is the stronger model overall, scoring 45.0 to 39.3 on the Noometry Index. Gemini 2.5 Flash costs 2.0× less per token, which makes it the better buy when GPT-5.4 mini's lead doesn't matter for your workload.

Last verified . 33 shared benchmarks.

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

GPT-5.4 mini OpenAI

45.0

Rank #76 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Gemini 2.5 Flash scores higher in 5 categories and GPT-5.4 mini in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.4 mini leads 51.5 to 36.4.
  • The biggest single-benchmark swing is ARC-AGI-1: 33.3% for Gemini 2.5 Flash and 63.7% for GPT-5.4 mini.
  • Gemini 2.5 Flash is cheaper at $0.30 / $2.50 per million input/output tokens, against $0.75 / $4.50 for GPT-5.4 mini.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 400K.

Side by side

Gemini 2.5 Flash and GPT-5.4 mini specifications
Gemini 2.5 FlashGPT-5.4 mini
ProviderGoogleOpenAI
Noometry Index39.345.0
Released2025-04-172026-03-17
WeightsProprietaryProprietary
Context window1.05M400K
Max output66K128K
Input $ / M tokens$0.30$0.75
Output $ / M tokens$2.50$4.50
Results tracked5446

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 mini leads

Gemini 2.5 Flash: 35.8 (#220), GPT-5.4 mini: 45.2 (#72)

Coding benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
WeirdML41.9%60.3%
LMArena Coding14241438
ALE-Bench661.881,189
FrontierCode—27%
SWE-bench Verified (bash only)28.7%—
Aider Polyglot55.1%—
LMArena WebDev—1397
SciCode—49.9%

Agentic & Tool Use Too close to call

Gemini 2.5 Flash: 30.8 (#74), GPT-5.4 mini: 29.9 (#81)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
Terminal-Bench17.1%—
Berkeley Function Calling Leaderboard56.2%—
TheAgentCompany41.1%—
DeepResearch Bench—36.3%
BALROG33.5%—
Vending-Bench 2548.84—

Reasoning GPT-5.4 mini leads

Gemini 2.5 Flash: 18.1 (#286), GPT-5.4 mini: 30.4 (#85)

Reasoning benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
ARC-AGI-22.5%18.9%
Kagi LLM Benchmark56.8%37.9%
ARC-AGI-133.3%63.7%
CritPt1.1%10%
LMArena Hard Prompts14221424
DTBench76.5%80%
LMCA27.5%40.8%
Epoch Capabilities Index143.03148.84
ForecastBench60.657
SimpleBench41.2%—
NYT Connections (extended)—61.8%
Chess Puzzles—24%
EnigmaEval2.7%—
Thematic Generalization—61.7%
Mystery Game Puzzles—11%

Math GPT-5.4 mini leads

Gemini 2.5 Flash: 39.9 (#98), GPT-5.4 mini: 45.5 (#75)

Math benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
OTIS Mock AIME 2024-202573.1%88.9%
LMArena Math14151419
FrontierMath (Feb 2025 set)4.8%28.3%
FrontierMath Tier 4 (v1)4.2%2.1%
FrontierMath (Tiers 1-3)—51.2%
FrontierMath Tier 4—9.8%
ProofBench—21%
Omni-MATH38.5%—

Knowledge GPT-5.4 mini leads

Gemini 2.5 Flash: 36.4 (#168), GPT-5.4 mini: 51.5 (#67)

Knowledge benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
Vectara Hallucination Rate7.8%5.5%
LMArena Expert14261435
GPQA Diamond—86.9%
Humanity's Last Exam12.1%—
SimpleQA Verified—29.4%
MMLU-Pro63.9%—
Confabulations16.8%—
GPQA (HELM)39%—

Multimodal Gemini 2.5 Flash leads

Gemini 2.5 Flash: 41.8 (#32), GPT-5.4 mini: 39.7 (#56)

Multimodal benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
LMArena Vision12531245
GeoBench76%—
VPCT46.2%—
SpatialViz-Bench36.9%—

Multilingual Too close to call

Gemini 2.5 Flash: 52.3 (#88), GPT-5.4 mini: 51.9 (#96)

Multilingual benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
LMArena Non-English14091405
LMArena Chinese14501446
LMArena French14331440
LMArena German14181409
LMArena Japanese14051374
LMArena Korean13851368
LMArena Russian14151417
LMArena Spanish14211405

Instruction Following Gemini 2.5 Flash leads

Gemini 2.5 Flash: 75.7 (#54), GPT-5.4 mini: 74.1 (#102)

Instruction Following benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
LMArena Instruction Following14051405
IFEval89.8%—

Long Context Gemini 2.5 Flash leads

Gemini 2.5 Flash: 47.5 (#17), GPT-5.4 mini: 43.0 (#112)

Long Context benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
LMArena Longer Query14191407
Fiction.LiveBench77.8%—

Writing & Preference GPT-5.4 mini leads

Gemini 2.5 Flash: 53.8 (#157), GPT-5.4 mini: 64.0 (#58)

Writing & Preference benchmarks
BenchmarkGemini 2.5 FlashGPT-5.4 mini
LMArena Text14171412
LMArena Creative Writing14001370
EQ-Bench Creative Writing11371665
LMArena Multi-Turn14081429
Short-Story Creative Writing76.5%—
WildBench81.7%—

Frequently asked questions

Is Gemini 2.5 Flash better than GPT-5.4 mini?

GPT-5.4 mini is the stronger model overall, scoring 45.0 to 39.3 on the Noometry Index. Gemini 2.5 Flash costs 2.0× less per token, which makes it the better buy when GPT-5.4 mini's lead doesn't matter for your workload.

Which is cheaper, Gemini 2.5 Flash or GPT-5.4 mini?

Gemini 2.5 Flash is cheaper. It lists at $0.30 per million input tokens and $2.50 per million output tokens; GPT-5.4 mini lists at $0.75 and $4.50.

Is Gemini 2.5 Flash or GPT-5.4 mini better for coding?

GPT-5.4 mini scores higher on coding benchmarks: 45.2 versus 35.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 400K.

How many benchmarks do Gemini 2.5 Flash and GPT-5.4 mini share?

33 benchmarks have published results for both models. Gemini 2.5 Flash has 54 scored results on Noometry and GPT-5.4 mini has 46.

Related comparisons

Go deeper