Model comparison

Gemini 3.8 Flash vs GPT-5.4

Gemini 3.8 Flash is the stronger model overall, scoring 61.8 to 59.4 on the Noometry Index.

Last verified . 44 shared benchmarks.

Gemini 3.8 Flash Google

61.8

Rank #11 Confirmed

GPT-5.4 OpenAI

59.4

Rank #16 Confirmed

Summary

  • They share 44 benchmarks with published results for both. Gemini 3.8 Flash scores higher in 6 categories and GPT-5.4 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.8 Flash leads 76.9 to 61.8.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 22% for Gemini 3.8 Flash and 49% for GPT-5.4.
  • Gemini 3.8 Flash is cheaper at $0.75 / $3.75 per million input/output tokens, against $2.50 / $15 for GPT-5.4.
  • GPT-5.4 accepts more context: 1.05M tokens versus 1.05M.

Side by side

Gemini 3.8 Flash and GPT-5.4 specifications
Gemini 3.8 FlashGPT-5.4
ProviderGoogleOpenAI
Noometry Index61.859.4
Released2026-09-022026-03-05
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output66K128K
Input $ / M tokens$0.75$2.50
Output $ / M tokens$3.75$15
Results tracked5068

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3.8 Flash leads

Gemini 3.8 Flash: 59.2 (#15), GPT-5.4: 52.6 (#33)

Coding benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
DeepSWE73.8%51.8%
LMArena WebDev15841465
SciCode56.6%56.6%
WeirdML84.8%77.7%
LMArena Coding15101497
ALE-Bench1,2701,607
SWE-bench Verified—76.9%
FrontierCode41.2%—
CursorBench39.6%—
FrontierSWE19.6%—
GSO—31.4%
MirrorCode—15.6%
AlgoTune—1.85

Agentic & Tool Use GPT-5.4 leads

Gemini 3.8 Flash: 41.8 (#21), GPT-5.4: 46.5 (#13)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
APEX-Agents64.3%52.4%
Vending-Bench 25,0946,144
Terminal-Bench—81.8%
Remote Labor Index5.8%—
τ²-bench Banking—39.4%
DeepResearch Bench—35.1%
PostTrainBench—19%
GBAEval—45.1%
GDP.pdf23.4%—
LMArena Search—1197
METR Time Horizons—74.3%

Reasoning Gemini 3.8 Flash leads

Gemini 3.8 Flash: 76.9 (#5), GPT-5.4: 61.8 (#19)

Reasoning benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
ARC-AGI-289.2%74%
NYT Connections (extended)97.4%91.3%
ARC-AGI-198.5%93.7%
CritPt18.3%23.4%
Chess Puzzles61%44%
LMArena Hard Prompts15081485
Mystery Game Puzzles47%37%
DTBench95.7%94.4%
LMCA52.9%52%
Epoch Capabilities Index156.71156.81
Kagi LLM Benchmark—63.8%
EnigmaEval—16%
Thematic Generalization—80%
EBR-Bench—25.4%
Surface Evolver Bench76.9%—
ForecastBench—59.5

Math GPT-5.4 leads

Gemini 3.8 Flash: 65.3 (#28), GPT-5.4: 73.5 (#19)

Knowledge Gemini 3.8 Flash leads

Gemini 3.8 Flash: 74.8 (#2), GPT-5.4: 65.3 (#14)

Knowledge benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
GPQA Diamond95.4%93.3%
Humanity's Last Exam44.5%36.2%
SimpleQA Verified69.7%45.1%
LMArena Expert15241507
Vectara Hallucination Rate—7%

Multimodal GPT-5.4 leads

Gemini 3.8 Flash: 40.7 (#45), GPT-5.4: 43.7 (#20)

Multimodal benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
LMArena Vision13141303
Blueprint-Bench 238.6%27.1%
Furniture Assembly31.7%37.5%
LMArena Document—1471

Multilingual Gemini 3.8 Flash leads

Gemini 3.8 Flash: 58.0 (#5), GPT-5.4: 56.2 (#23)

Multilingual benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
LMArena Non-English14911465
LMArena Chinese15541519
LMArena French14981493
LMArena German14931472
LMArena Japanese15021485
LMArena Korean14591448
LMArena Russian15151480
LMArena Spanish14851454

Instruction Following Too close to call

Gemini 3.8 Flash: 78.0 (#13), GPT-5.4: 77.1 (#27)

Instruction Following benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
LMArena Instruction Following14901469

Long Context GPT-5.4 leads

Gemini 3.8 Flash: 46.3 (#24), GPT-5.4: 50.3 (#8)

Long Context benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
LMArena Longer Query15081473
CL-bench—27.9%
CL-bench Life—21.7%

Writing & Preference Too close to call

Gemini 3.8 Flash: 72.2 (#15), GPT-5.4: 71.9 (#17)

Writing & Preference benchmarks
BenchmarkGemini 3.8 FlashGPT-5.4
LMArena Text14991469
LMArena Creative Writing14921439
EQ-Bench Creative Writing17481840
LMArena Multi-Turn15011482
EQ-Bench 4—1272

Frequently asked questions

Is Gemini 3.8 Flash better than GPT-5.4?

Gemini 3.8 Flash is the stronger model overall, scoring 61.8 to 59.4 on the Noometry Index.

Which is cheaper, Gemini 3.8 Flash or GPT-5.4?

Gemini 3.8 Flash is cheaper. It lists at $0.75 per million input tokens and $3.75 per million output tokens; GPT-5.4 lists at $2.50 and $15.

Is Gemini 3.8 Flash or GPT-5.4 better for coding?

Gemini 3.8 Flash scores higher on coding benchmarks: 59.2 versus 52.6 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 does, with 1.05M tokens against 1.05M.

How many benchmarks do Gemini 3.8 Flash and GPT-5.4 share?

44 benchmarks have published results for both models. Gemini 3.8 Flash has 50 scored results on Noometry and GPT-5.4 has 68.

Related comparisons

Go deeper