Model comparison

Gemini 3 Pro vs GPT-5.5

GPT-5.5 is the stronger model overall, scoring 63.4 to 54.8 on the Noometry Index.

Last verified . 47 shared benchmarks.

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

GPT-5.5 OpenAI

63.4

Rank #9 Confirmed

Summary

  • They share 47 benchmarks with published results for both. Gemini 3 Pro scores higher in 3 categories and GPT-5.5 in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.5 leads 81.7 to 49.9.
  • The biggest single-benchmark swing is ARC-AGI-2: 31.1% for Gemini 3 Pro and 85% for GPT-5.5.

Side by side

Gemini 3 Pro and GPT-5.5 specifications
Gemini 3 ProGPT-5.5
ProviderGoogleOpenAI
Noometry Index54.863.4
Released2025-11-182026-04-23
WeightsProprietaryProprietary
Context window—1.05M
Max output—128K
Input $ / M tokens—$5
Output $ / M tokens—$30
Results tracked6771

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 leads

Gemini 3 Pro: 51.6 (#39), GPT-5.5: 58.2 (#17)

Coding benchmarks
BenchmarkGemini 3 ProGPT-5.5
SWE-bench Verified72.9%80.6%
LMArena WebDev14401513
GSO18.6%40.2%
WeirdML69.9%84.9%
LMArena Coding14811494
ALE-Bench1,1771,943
DeepSWE—67%
FrontierCode—43%
SWE-bench Verified (bash only)74.2%—
SWE-bench Multilingual68.7%—
SciCode—56.1%
MirrorCode—10%
AlgoTune1.83—

Agentic & Tool Use GPT-5.5 leads

Gemini 3 Pro: 40.6 (#23), GPT-5.5: 50.7 (#6)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 ProGPT-5.5
Terminal-Bench69.4%84.7%
Remote Labor Index1.3%6.3%
τ²-bench Banking18%44.6%
DeepResearch Bench46.3%54%
LMArena Search12071242
Vending-Bench 25,4787,524
APEX-Agents—55.1%
Berkeley Function Calling Leaderboard72.5%—
OSWorld 2.0—13%
GDPval40.3%—
τ²-bench Airline80.5%—
τ²-bench Retail75.9%—
τ²-bench Telecom91%—
PostTrainBench—27.2%
BALROG58.1%—
ExploitBench—47.4%
GBAEval—53.2%
GDP.pdf—26%
METR Time Horizons71%—

Reasoning GPT-5.5 leads

Gemini 3 Pro: 52.5 (#31), GPT-5.5: 72.8 (#11)

Reasoning benchmarks
BenchmarkGemini 3 ProGPT-5.5
ARC-AGI-231.1%85%
SimpleBench76.4%69%
Kagi LLM Benchmark80.1%88.8%
NYT Connections (extended)94.4%96.2%
ARC-AGI-175%95%
CritPt6.9%27.1%
Chess Puzzles31%54%
LMArena Hard Prompts14801489
Epoch Capabilities Index152.92159.1
ForecastBench61.260.6
EnigmaEval18.2%—
EBR-Bench—34.3%
Mystery Game Puzzles—56%
DTBench—96%
LMCA—54.3%
Surface Evolver Bench—88.1%
Bench to the Future 3—0.14

Math GPT-5.5 leads

Gemini 3 Pro: 49.9 (#59), GPT-5.5: 81.7 (#11)

Knowledge Too close to call

Gemini 3 Pro: 64.4 (#16), GPT-5.5: 64.4 (#17)

Knowledge benchmarks
BenchmarkGemini 3 ProGPT-5.5
GPQA Diamond92.6%94%
Vectara Hallucination Rate13.6%9.3%
LMArena Expert14751508
Humanity's Last Exam37.5%—
SimpleQA Verified—63%
MMLU-Pro90.3%—
GPQA (HELM)80.3%—

Multimodal Gemini 3 Pro leads

Gemini 3 Pro: 57.6 (#2), GPT-5.5: 46.9 (#12)

Multimodal benchmarks
BenchmarkGemini 3 ProGPT-5.5
LMArena Vision13051297
LMArena Document14341486
GeoBench84%—
VPCT91%—
Blueprint-Bench 2—36.2%
Furniture Assembly—44.2%

Multilingual Too close to call

Gemini 3 Pro: 56.9 (#16), GPT-5.5: 56.4 (#20)

Multilingual benchmarks
BenchmarkGemini 3 ProGPT-5.5
LMArena Non-English14741467
LMArena Chinese15231533
LMArena French14921486
LMArena German15151480
LMArena Japanese15101498
LMArena Korean14481460
LMArena Russian14931473
LMArena Spanish14701468

Instruction Following GPT-5.5 leads

Gemini 3 Pro: 76.3 (#45), GPT-5.5: 77.5 (#18)

Instruction Following benchmarks
BenchmarkGemini 3 ProGPT-5.5
LMArena Instruction Following14581479
IFEval87.7%—

Long Context GPT-5.5 leads

Gemini 3 Pro: 44.0 (#79), GPT-5.5: 48.3 (#12)

Long Context benchmarks
BenchmarkGemini 3 ProGPT-5.5
LMArena Longer Query14711484
CL-bench15.8%—
CL-bench Life—22.2%

Writing & Preference GPT-5.5 leads

Gemini 3 Pro: 66.4 (#35), GPT-5.5: 72.7 (#13)

Writing & Preference benchmarks
BenchmarkGemini 3 ProGPT-5.5
LMArena Text14791472
LMArena Creative Writing14821455
EQ-Bench Creative Writing15251844
LMArena Multi-Turn14841476
WildBench85.9%—
EQ-Bench 4—1315

Frequently asked questions

Is Gemini 3 Pro better than GPT-5.5?

GPT-5.5 is the stronger model overall, scoring 63.4 to 54.8 on the Noometry Index.

Is Gemini 3 Pro or GPT-5.5 better for coding?

GPT-5.5 scores higher on coding benchmarks: 58.2 versus 51.6 in the Noometry coding category.

How many benchmarks do Gemini 3 Pro and GPT-5.5 share?

47 benchmarks have published results for both models. Gemini 3 Pro has 67 scored results on Noometry and GPT-5.5 has 71.

Related comparisons

Go deeper