Model comparison

Gemini 1.5 Flash 8B vs gpt-oss-120b

gpt-oss-120b is the stronger model overall, scoring 36.3 to 29.9 on the Noometry Index.

Last verified . 20 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 3 categories and gpt-oss-120b in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-120b leads 52.5 to 14.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.6% for Gemini 1.5 Flash 8B and 88.9% for gpt-oss-120b.
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash 8B and gpt-oss-120b specifications
Gemini 1.5 Flash 8Bgpt-oss-120b
ProviderGoogleOpenAI
Noometry Index29.936.3
Released2024-10-032025-08-05
WeightsProprietaryOpen
Context window—131K
Max output—41K
Input $ / M tokens—$0.037
Output $ / M tokens—$0.17
Results tracked2148

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 35.5 (#225), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Coding12181380
SWE-bench Verified (bash only)—26%
Aider Polyglot—41.8%
SciCode—36%
WeirdML—48.2%
ALE-Bench—575.62
AlgoTune—1.41

Agentic & Tool Use Not comparable

Gemini 1.5 Flash 8B: —, gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
Terminal-Bench—18.7%
APEX-Agents—4.4%
METR Time Horizons—56.6%
Vending-Bench 2—-21.53

Reasoning Too close to call

Gemini 1.5 Flash 8B: 20.0 (#244), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Hard Prompts12091364
DTBench50%76.3%
SimpleBench—22.1%
Kagi LLM Benchmark—58.6%
CritPt—1.1%
Chess Puzzles—20%
Mystery Game Puzzles—2%
LMCA—22.1%
Surface Evolver Bench—25%
Epoch Capabilities Index—139.93

Math gpt-oss-120b leads

Gemini 1.5 Flash 8B: 14.2 (#302), gpt-oss-120b: 52.5 (#50)

Math benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
OTIS Mock AIME 2024-20254.6%88.9%
LMArena Math12071389
Omni-MATH—68.8%

Knowledge gpt-oss-120b leads

Gemini 1.5 Flash 8B: 16.0 (#289), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
GPQA Diamond33%75.8%
LMArena Expert11851356
MMLU-Pro—79.5%
Confabulations—15.7%
Vectara Hallucination Rate—14.2%
GPQA (HELM)—68.4%

Multimodal Not comparable

Gemini 1.5 Flash 8B: 28.2 (#115), gpt-oss-120b: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Vision1044—

Multilingual gpt-oss-120b leads

Gemini 1.5 Flash 8B: 38.5 (#229), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Non-English12151351
LMArena Chinese12311385
LMArena French12341369
LMArena German12061353
LMArena Japanese11501331
LMArena Korean11401282
LMArena Russian12361343
LMArena Spanish12121389

Instruction Following gpt-oss-120b leads

Gemini 1.5 Flash 8B: 62.8 (#236), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Instruction Following11991318
IFEval—83.6%

Long Context Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 37.0 (#225), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Longer Query12191319
Fiction.LiveBench—44.4%

Writing & Preference gpt-oss-120b leads

Gemini 1.5 Flash 8B: 42.8 (#232), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8Bgpt-oss-120b
LMArena Text12261365
LMArena Creative Writing12181275
LMArena Multi-Turn11851340
Short-Story Creative Writing—77.1%
EQ-Bench Creative Writing—961
WildBench—84.5%

Frequently asked questions

Is Gemini 1.5 Flash 8B better than gpt-oss-120b?

gpt-oss-120b is the stronger model overall, scoring 36.3 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or gpt-oss-120b better for coding?

Gemini 1.5 Flash 8B scores higher on coding benchmarks: 35.5 versus 33.5 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and gpt-oss-120b share?

20 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper