Model comparison

Gemini 1.5 Flash (May 2024) vs GPT-4.5

GPT-4.5 is the stronger model overall, scoring 37.2 to 33.2 on the Noometry Index.

Last verified . 23 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 1 category and GPT-4.5 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-4.5 leads 32.6 to 22.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 16.3% for Gemini 1.5 Flash (May 2024) and 37.8% for GPT-4.5.

Side by side

Gemini 1.5 Flash (May 2024) and GPT-4.5 specifications
Gemini 1.5 Flash (May 2024)GPT-4.5
ProviderGoogleOpenAI
Noometry Index33.237.2
Released2024-05-142025-02-27
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), GPT-4.5: 42.2 (#109)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
WeirdML24.9%39.4%
LMArena Coding12611396
Aider Polyglot—44.9%
BigCodeBench Instruct43.5%—
LiveBench Coding—75.2%
BigCodeBench Complete55.1%—
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 26.6 (#102), GPT-4.5: 27.9 (#97)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
Cybench—17.5%
BALROG14.6%—

Reasoning Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), GPT-4.5: 13.9 (#330)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
LMArena Hard Prompts12571403
Epoch Capabilities Index129.36136.74
ForecastBench53.961.7
ARC-AGI-2—0.8%
SimpleBench—34.5%
ARC-AGI-1—10.3%
EnigmaEval—3.2%
LiveBench Reasoning—71.1%
DTBench53.8%—
LiveBench Data Analysis—64.3%
LiveBench—69%
PIQA87.5%—

Math GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), GPT-4.5: 32.6 (#211)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
OTIS Mock AIME 2024-202516.3%37.8%
LMArena Math12691412
MATH Level 561.9%78.6%
Omni-MATH30.4%—
LiveBench Math—69.3%
FrontierMath (Feb 2025 set)0%—
GSM8K82.4%—

Knowledge GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), GPT-4.5: 32.5 (#211)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
GPQA Diamond47.3%68.7%
LMArena Expert12331394
Humanity's Last Exam—5.4%
MMLU-Pro67.8%—
Confabulations—13.6%
GPQA (HELM)43.7%—
BoolQ85.8%—
MMLU77.9%—

Multimodal GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 36.0 (#81), GPT-4.5: 37.6 (#71)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
LMArena Vision11411195
Video-MME70.3%—
GeoBench76%—
VPCT—45%

Multilingual GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), GPT-4.5: 52.5 (#83)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
LMArena Non-English12781413
LMArena Chinese12951421
LMArena French12581418
LMArena German12621457
LMArena Japanese12521416
LMArena Korean12211392
LMArena Russian12881419
LMArena Spanish1243—

Instruction Following GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), GPT-4.5: 72.6 (#134)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
LMArena Instruction Following12581404
LiveBench Instruction Following—72.3%
IFEval83.1%—

Long Context GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), GPT-4.5: 40.4 (#155)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
LMArena Longer Query12841406
Fiction.LiveBench—63.9%

Writing & Preference GPT-4.5 leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), GPT-4.5: 56.9 (#134)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4.5
LMArena Text12871417
LMArena Creative Writing12851394
LMArena Multi-Turn12531444
Short-Story Creative Writing—75.6%
EQ-Bench Creative Writing—1258
WildBench79.2%—
LiveBench Language—61.5%

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than GPT-4.5?

GPT-4.5 is the stronger model overall, scoring 37.2 to 33.2 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or GPT-4.5 better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 34.4 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and GPT-4.5 share?

23 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and GPT-4.5 has 42.

Related comparisons

Go deeper