Model comparison

Gemini 3 Pro vs gpt-oss-120b

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 36.3 on the Noometry Index.

Last verified . 38 shared benchmarks.

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Gemini 3 Pro scores higher in 8 categories and gpt-oss-120b in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3 Pro leads 52.5 to 20.0.
  • The biggest single-benchmark swing is SimpleBench: 76.4% for Gemini 3 Pro and 22.1% for gpt-oss-120b.
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

Gemini 3 Pro and gpt-oss-120b specifications
Gemini 3 Progpt-oss-120b
ProviderGoogleOpenAI
Noometry Index54.836.3
Released2025-11-182025-08-05
WeightsProprietaryOpen
Context window—131K
Max output—41K
Input $ / M tokens—$0.037
Output $ / M tokens—$0.17
Results tracked6748

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3 Pro leads

Gemini 3 Pro: 51.6 (#39), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkGemini 3 Progpt-oss-120b
SWE-bench Verified (bash only)74.2%26%
WeirdML69.9%48.2%
LMArena Coding14811380
ALE-Bench1,177575.62
AlgoTune1.831.41
SWE-bench Verified72.9%—
Aider Polyglot—41.8%
LMArena WebDev1440—
SWE-bench Multilingual68.7%—
SciCode—36%
GSO18.6%—

Agentic & Tool Use Gemini 3 Pro leads

Gemini 3 Pro: 40.6 (#23), gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Progpt-oss-120b
Terminal-Bench69.4%18.7%
METR Time Horizons71%56.6%
Vending-Bench 25,478-21.53
APEX-Agents—4.4%
Berkeley Function Calling Leaderboard72.5%—
GDPval40.3%—
Remote Labor Index1.3%—
τ²-bench Airline80.5%—
τ²-bench Banking18%—
τ²-bench Retail75.9%—
τ²-bench Telecom91%—
DeepResearch Bench46.3%—
BALROG58.1%—
LMArena Search1207—

Reasoning Gemini 3 Pro leads

Gemini 3 Pro: 52.5 (#31), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkGemini 3 Progpt-oss-120b
SimpleBench76.4%22.1%
Kagi LLM Benchmark80.1%58.6%
CritPt6.9%1.1%
Chess Puzzles31%20%
LMArena Hard Prompts14801364
Epoch Capabilities Index152.92139.93
ARC-AGI-231.1%—
NYT Connections (extended)94.4%—
ARC-AGI-175%—
EnigmaEval18.2%—
Mystery Game Puzzles—2%
DTBench—76.3%
LMCA—22.1%
Surface Evolver Bench—25%
ForecastBench61.2—

Math gpt-oss-120b leads

Gemini 3 Pro: 49.9 (#59), gpt-oss-120b: 52.5 (#50)

Math benchmarks
BenchmarkGemini 3 Progpt-oss-120b
OTIS Mock AIME 2024-202591.4%88.9%
Omni-MATH55.5%68.8%
LMArena Math14761389
MathArena Final-Answer Competitions67%—
ProofBench20%—
FrontierMath (Feb 2025 set)37.6%—
FrontierMath Tier 4 (v1)18.8%—

Knowledge Gemini 3 Pro leads

Gemini 3 Pro: 64.4 (#16), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkGemini 3 Progpt-oss-120b
GPQA Diamond92.6%75.8%
MMLU-Pro90.3%79.5%
Vectara Hallucination Rate13.6%14.2%
GPQA (HELM)80.3%68.4%
LMArena Expert14751356
Humanity's Last Exam37.5%—
Confabulations—15.7%

Multimodal Not comparable

Gemini 3 Pro: 57.6 (#2), gpt-oss-120b: —

Multimodal benchmarks
BenchmarkGemini 3 Progpt-oss-120b
LMArena Vision1305—
GeoBench84%—
VPCT91%—
LMArena Document1434—

Multilingual Gemini 3 Pro leads

Gemini 3 Pro: 56.9 (#16), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkGemini 3 Progpt-oss-120b
LMArena Non-English14741351
LMArena Chinese15231385
LMArena French14921369
LMArena German15151353
LMArena Japanese15101331
LMArena Korean14481282
LMArena Russian14931343
LMArena Spanish14701389

Instruction Following Gemini 3 Pro leads

Gemini 3 Pro: 76.3 (#45), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkGemini 3 Progpt-oss-120b
IFEval87.7%83.6%
LMArena Instruction Following14581318

Long Context Gemini 3 Pro leads

Gemini 3 Pro: 44.0 (#79), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkGemini 3 Progpt-oss-120b
LMArena Longer Query14711319
Fiction.LiveBench—44.4%
CL-bench15.8%—

Writing & Preference Gemini 3 Pro leads

Gemini 3 Pro: 66.4 (#35), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkGemini 3 Progpt-oss-120b
LMArena Text14791365
LMArena Creative Writing14821275
EQ-Bench Creative Writing1525961
WildBench85.9%84.5%
LMArena Multi-Turn14841340
Short-Story Creative Writing—77.1%

Frequently asked questions

Is Gemini 3 Pro better than gpt-oss-120b?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 36.3 on the Noometry Index.

Is Gemini 3 Pro or gpt-oss-120b better for coding?

Gemini 3 Pro scores higher on coding benchmarks: 51.6 versus 33.5 in the Noometry coding category.

How many benchmarks do Gemini 3 Pro and gpt-oss-120b share?

38 benchmarks have published results for both models. Gemini 3 Pro has 67 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper