Model comparison

Gemini 1.5 Flash (May 2024) vs Gemini 3 Pro

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 33.2 on the Noometry Index.

Last verified . 31 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 0 categories and Gemini 3 Pro in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemini 3 Pro leads 64.4 to 26.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 16.3% for Gemini 1.5 Flash (May 2024) and 91.4% for Gemini 3 Pro.

Side by side

Gemini 1.5 Flash (May 2024) and Gemini 3 Pro specifications
Gemini 1.5 Flash (May 2024)Gemini 3 Pro
ProviderGoogleGoogle
Noometry Index33.254.8
Released2024-05-142025-11-18
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4267

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), Gemini 3 Pro: 51.6 (#39)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
WeirdML24.9%69.9%
LMArena Coding12611481
SWE-bench Verified—72.9%
SWE-bench Verified (bash only)—74.2%
LMArena WebDev—1440
SWE-bench Multilingual—68.7%
GSO—18.6%
BigCodeBench Instruct43.5%—
BigCodeBench Complete55.1%—
ALE-Bench—1,177
AlgoTune—1.83
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 26.6 (#102), Gemini 3 Pro: 40.6 (#23)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
BALROG14.6%58.1%
Terminal-Bench—69.4%
Berkeley Function Calling Leaderboard—72.5%
GDPval—40.3%
Remote Labor Index—1.3%
τ²-bench Airline—80.5%
τ²-bench Banking—18%
τ²-bench Retail—75.9%
τ²-bench Telecom—91%
DeepResearch Bench—46.3%
LMArena Search—1207
METR Time Horizons—71%
Vending-Bench 2—5,478

Reasoning Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), Gemini 3 Pro: 52.5 (#31)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
LMArena Hard Prompts12571480
Epoch Capabilities Index129.36152.92
ForecastBench53.961.2
ARC-AGI-2—31.1%
SimpleBench—76.4%
Kagi LLM Benchmark—80.1%
NYT Connections (extended)—94.4%
ARC-AGI-1—75%
CritPt—6.9%
Chess Puzzles—31%
EnigmaEval—18.2%
DTBench53.8%—
PIQA87.5%—

Math Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Gemini 3 Pro: 49.9 (#59)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
OTIS Mock AIME 2024-202516.3%91.4%
Omni-MATH30.4%55.5%
LMArena Math12691476
FrontierMath (Feb 2025 set)0%37.6%
MathArena Final-Answer Competitions—67%
ProofBench—20%
MATH Level 561.9%—
FrontierMath Tier 4 (v1)—18.8%
GSM8K82.4%—

Knowledge Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Gemini 3 Pro: 64.4 (#16)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
GPQA Diamond47.3%92.6%
MMLU-Pro67.8%90.3%
GPQA (HELM)43.7%80.3%
LMArena Expert12331475
Humanity's Last Exam—37.5%
Vectara Hallucination Rate—13.6%
BoolQ85.8%—
MMLU77.9%—

Multimodal Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 36.0 (#81), Gemini 3 Pro: 57.6 (#2)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
LMArena Vision11411305
GeoBench76%84%
Video-MME70.3%—
VPCT—91%
LMArena Document—1434

Multilingual Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Gemini 3 Pro: 56.9 (#16)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
LMArena Non-English12781474
LMArena Chinese12951523
LMArena French12581492
LMArena German12621515
LMArena Japanese12521510
LMArena Korean12211448
LMArena Russian12881493
LMArena Spanish12431470

Instruction Following Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Gemini 3 Pro: 76.3 (#45)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
IFEval83.1%87.7%
LMArena Instruction Following12581458

Long Context Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), Gemini 3 Pro: 44.0 (#79)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
LMArena Longer Query12841471
CL-bench—15.8%

Writing & Preference Gemini 3 Pro leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Gemini 3 Pro: 66.4 (#35)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Gemini 3 Pro
LMArena Text12871479
LMArena Creative Writing12851482
WildBench79.2%85.9%
LMArena Multi-Turn12531484
EQ-Bench Creative Writing—1525

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Gemini 3 Pro?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 33.2 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Gemini 3 Pro better for coding?

Gemini 3 Pro scores higher on coding benchmarks: 51.6 versus 34.4 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Gemini 3 Pro share?

31 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Gemini 3 Pro has 67.

Related comparisons

Go deeper