Model comparison

Gemini 3 Pro vs GPT-4o

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 28.6 on the Noometry Index.

Last verified . 46 shared benchmarks.

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

GPT-4o OpenAI

28.6

Rank #324 Confirmed

Summary

  • They share 46 benchmarks with published results for both. Gemini 3 Pro scores higher in 10 categories and GPT-4o in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3 Pro leads 52.5 to 9.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 91.4% for Gemini 3 Pro and 6.4% for GPT-4o.

Side by side

Gemini 3 Pro and GPT-4o specifications
Gemini 3 ProGPT-4o
ProviderGoogleOpenAI
Noometry Index54.828.6
Released2025-11-182024-05-13
WeightsProprietaryProprietary
Context window—128K
Max output—16K
Input $ / M tokens—$2.50
Output $ / M tokens—$10
Results tracked6772

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3 Pro leads

Gemini 3 Pro: 51.6 (#39), GPT-4o: 24.8 (#328)

Coding benchmarks
BenchmarkGemini 3 ProGPT-4o
SWE-bench Verified72.9%31%
SWE-bench Verified (bash only)74.2%21.6%
GSO18.6%0%
WeirdML69.9%25.1%
LMArena Coding14811297
Aider Polyglot—45.3%
LMArena WebDev1440—
SWE-bench Multilingual68.7%—
BigCodeBench Instruct—51.1%
LiveBench Coding—51.4%
BigCodeBench Complete—61.1%
CadEval—26%
ALE-Bench1,177—
AlgoTune1.83—
HumanEval+—87.2%
MBPP+—72.2%

Agentic & Tool Use Gemini 3 Pro leads

Gemini 3 Pro: 40.6 (#23), GPT-4o: 21.0 (#141)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 ProGPT-4o
GDPval40.3%9.9%
BALROG58.1%32.3%
LMArena Search12071006
METR Time Horizons71%40.8%
Terminal-Bench69.4%—
Berkeley Function Calling Leaderboard72.5%—
Remote Labor Index1.3%—
TheAgentCompany—8.6%
τ²-bench Airline80.5%—
τ²-bench Banking18%—
τ²-bench Retail75.9%—
τ²-bench Telecom91%—
Cybench—12.5%
DeepResearch Bench46.3%—
Vending-Bench 25,478—

Reasoning Gemini 3 Pro leads

Gemini 3 Pro: 52.5 (#31), GPT-4o: 9.4 (#343)

Reasoning benchmarks
BenchmarkGemini 3 ProGPT-4o
ARC-AGI-231.1%0%
SimpleBench76.4%17.8%
ARC-AGI-175%4.5%
CritPt6.9%0%
Chess Puzzles31%13%
EnigmaEval18.2%0.8%
LMArena Hard Prompts14801281
Epoch Capabilities Index152.92128.97
ForecastBench61.257.7
Kagi LLM Benchmark80.1%—
NYT Connections (extended)94.4%—
LiveBench Reasoning—55.8%
DTBench—64.5%
LiveBench Data Analysis—60.9%
LMCA—16.6%
LiveBench—55.3%

Math Gemini 3 Pro leads

Gemini 3 Pro: 49.9 (#59), GPT-4o: 10.6 (#312)

Knowledge Gemini 3 Pro leads

Gemini 3 Pro: 64.4 (#16), GPT-4o: 28.8 (#242)

Knowledge benchmarks
BenchmarkGemini 3 ProGPT-4o
GPQA Diamond92.6%49.2%
Humanity's Last Exam37.5%2.7%
MMLU-Pro90.3%71.3%
Vectara Hallucination Rate13.6%9.6%
GPQA (HELM)80.3%52%
LMArena Expert14751250
SimpleQA Verified—26%
Confabulations—15.3%
MMLU—88.1%

Multimodal Gemini 3 Pro leads

Gemini 3 Pro: 57.6 (#2), GPT-4o: 34.5 (#91)

Multimodal benchmarks
BenchmarkGemini 3 ProGPT-4o
LMArena Vision13051137
GeoBench84%71%
VPCT91%40%
Video-MME—71.9%
LMArena Document1434—
ScienceQA—88.5%

Multilingual Gemini 3 Pro leads

Gemini 3 Pro: 56.9 (#16), GPT-4o: 43.2 (#186)

Multilingual benchmarks
BenchmarkGemini 3 ProGPT-4o
LMArena Non-English14741283
LMArena Chinese15231277
LMArena French14921304
LMArena German15151282
LMArena Japanese15101257
LMArena Korean14481234
LMArena Russian14931286
LMArena Spanish14701292

Instruction Following Gemini 3 Pro leads

Gemini 3 Pro: 76.3 (#45), GPT-4o: 66.6 (#207)

Instruction Following benchmarks
BenchmarkGemini 3 ProGPT-4o
IFEval87.7%81.7%
LMArena Instruction Following14581278
LiveBench Instruction Following—68.6%

Long Context Gemini 3 Pro leads

Gemini 3 Pro: 44.0 (#79), GPT-4o: 39.4 (#179)

Long Context benchmarks
BenchmarkGemini 3 ProGPT-4o
LMArena Longer Query14711289
Fiction.LiveBench—66.7%
CL-bench15.8%—

Writing & Preference Gemini 3 Pro leads

Gemini 3 Pro: 66.4 (#35), GPT-4o: 52.6 (#166)

Writing & Preference benchmarks
BenchmarkGemini 3 ProGPT-4o
LMArena Text14791300
LMArena Creative Writing14821292
WildBench85.9%82.8%
LMArena Multi-Turn14841302
Short-Story Creative Writing—81.8%
EQ-Bench Creative Writing1525—
LiveBench Language—47.6%

Frequently asked questions

Is Gemini 3 Pro better than GPT-4o?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 28.6 on the Noometry Index.

Is Gemini 3 Pro or GPT-4o better for coding?

Gemini 3 Pro scores higher on coding benchmarks: 51.6 versus 24.8 in the Noometry coding category.

How many benchmarks do Gemini 3 Pro and GPT-4o share?

46 benchmarks have published results for both models. Gemini 3 Pro has 67 scored results on Noometry and GPT-4o has 72.

Related comparisons

Go deeper