Model comparison

Gemini 2.5 Flash vs gpt-oss-120b

Gemini 2.5 Flash is the stronger model overall, scoring 39.3 to 36.3 on the Noometry Index. gpt-oss-120b costs 12× less per token, which makes it the better buy when Gemini 2.5 Flash's lead doesn't matter for your workload.

Last verified . 40 shared benchmarks.

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 40 benchmarks with published results for both. Gemini 2.5 Flash scores higher in 6 categories and gpt-oss-120b in 3 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Gemini 2.5 Flash leads 30.8 to 12.2.
  • The biggest single-benchmark swing is Fiction.LiveBench: 77.8% for Gemini 2.5 Flash and 44.4% for gpt-oss-120b.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $0.30 / $2.50 for Gemini 2.5 Flash.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 131K.
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

Gemini 2.5 Flash and gpt-oss-120b specifications
Gemini 2.5 Flashgpt-oss-120b
ProviderGoogleOpenAI
Noometry Index39.336.3
Released2025-04-172025-08-05
WeightsProprietaryOpen
Context window1.05M131K
Max output66K41K
Input $ / M tokens$0.30$0.037
Output $ / M tokens$2.50$0.17
Results tracked5448

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.5 Flash leads

Gemini 2.5 Flash: 35.8 (#220), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
SWE-bench Verified (bash only)28.7%26%
Aider Polyglot55.1%41.8%
WeirdML41.9%48.2%
LMArena Coding14241380
ALE-Bench661.88575.62
SciCode—36%
AlgoTune—1.41

Agentic & Tool Use Gemini 2.5 Flash leads

Gemini 2.5 Flash: 30.8 (#74), gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
Terminal-Bench17.1%18.7%
Vending-Bench 2548.84-21.53
APEX-Agents—4.4%
Berkeley Function Calling Leaderboard56.2%—
TheAgentCompany41.1%—
BALROG33.5%—
METR Time Horizons—56.6%

Reasoning gpt-oss-120b leads

Gemini 2.5 Flash: 18.1 (#286), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
SimpleBench41.2%22.1%
Kagi LLM Benchmark56.8%58.6%
CritPt1.1%1.1%
LMArena Hard Prompts14221364
DTBench76.5%76.3%
LMCA27.5%22.1%
Epoch Capabilities Index143.03139.93
ARC-AGI-22.5%—
ARC-AGI-133.3%—
Chess Puzzles—20%
EnigmaEval2.7%—
Mystery Game Puzzles—2%
Surface Evolver Bench—25%
ForecastBench60.6—

Math gpt-oss-120b leads

Gemini 2.5 Flash: 39.9 (#98), gpt-oss-120b: 52.5 (#50)

Math benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
OTIS Mock AIME 2024-202573.1%88.9%
Omni-MATH38.5%68.8%
LMArena Math14151389
FrontierMath (Feb 2025 set)4.8%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge gpt-oss-120b leads

Gemini 2.5 Flash: 36.4 (#168), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
MMLU-Pro63.9%79.5%
Confabulations16.8%15.7%
Vectara Hallucination Rate7.8%14.2%
GPQA (HELM)39%68.4%
LMArena Expert14261356
GPQA Diamond—75.8%
Humanity's Last Exam12.1%—

Multimodal Not comparable

Gemini 2.5 Flash: 41.8 (#32), gpt-oss-120b: —

Multimodal benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
LMArena Vision1253—
GeoBench76%—
VPCT46.2%—
SpatialViz-Bench36.9%—

Multilingual Gemini 2.5 Flash leads

Gemini 2.5 Flash: 52.3 (#88), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
LMArena Non-English14091351
LMArena Chinese14501385
LMArena French14331369
LMArena German14181353
LMArena Japanese14051331
LMArena Korean13851282
LMArena Russian14151343
LMArena Spanish14211389

Instruction Following Gemini 2.5 Flash leads

Gemini 2.5 Flash: 75.7 (#54), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
IFEval89.8%83.6%
LMArena Instruction Following14051318

Long Context Gemini 2.5 Flash leads

Gemini 2.5 Flash: 47.5 (#17), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
Fiction.LiveBench77.8%44.4%
LMArena Longer Query14191319

Writing & Preference Gemini 2.5 Flash leads

Gemini 2.5 Flash: 53.8 (#157), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flashgpt-oss-120b
LMArena Text14171365
LMArena Creative Writing14001275
Short-Story Creative Writing76.5%77.1%
EQ-Bench Creative Writing1137961
WildBench81.7%84.5%
LMArena Multi-Turn14081340

Frequently asked questions

Is Gemini 2.5 Flash better than gpt-oss-120b?

Gemini 2.5 Flash is the stronger model overall, scoring 39.3 to 36.3 on the Noometry Index. gpt-oss-120b costs 12× less per token, which makes it the better buy when Gemini 2.5 Flash's lead doesn't matter for your workload.

Which is cheaper, Gemini 2.5 Flash or gpt-oss-120b?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Gemini 2.5 Flash lists at $0.30 and $2.50.

Is Gemini 2.5 Flash or gpt-oss-120b better for coding?

Gemini 2.5 Flash scores higher on coding benchmarks: 35.8 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 131K.

How many benchmarks do Gemini 2.5 Flash and gpt-oss-120b share?

40 benchmarks have published results for both models. Gemini 2.5 Flash has 54 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper