Model comparison

Gemini 2.5 Flash-Lite vs gpt-oss-120b

Gemini 2.5 Flash-Lite and gpt-oss-120b score almost the same on the Noometry Index (37.0 vs 36.3), so choose on price, context window or the category you care about most.

Last verified . 30 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 7 categories and gpt-oss-120b in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Gemini 2.5 Flash-Lite leads 28.0 to 12.2.
  • The biggest single-benchmark swing is GPQA (HELM): 30.9% for Gemini 2.5 Flash-Lite and 68.4% for gpt-oss-120b.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $0.10 / $0.40 for Gemini 2.5 Flash-Lite.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 131K.
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

Gemini 2.5 Flash-Lite and gpt-oss-120b specifications
Gemini 2.5 Flash-Litegpt-oss-120b
ProviderGoogleOpenAI
Noometry Index37.036.3
Released2025-06-172025-08-05
WeightsProprietaryOpen
Context window1.05M131K
Max output66K41K
Input $ / M tokens$0.10$0.037
Output $ / M tokens$0.40$0.17
Results tracked3348

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.5 (#173), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
WeirdML35.2%48.2%
LMArena Coding13731380
ALE-Bench325.9575.62
SWE-bench Verified (bash only)—26%
Aider Polyglot—41.8%
SciCode—36%
AlgoTune—1.41

Agentic & Tool Use Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 28.0 (#96), gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
Terminal-Bench—18.7%
APEX-Agents—4.4%
Berkeley Function Calling Leaderboard36.9%—
METR Time Horizons—56.6%
Vending-Bench 2—-21.53

Reasoning Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 22.2 (#205), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
Kagi LLM Benchmark40.5%58.6%
LMArena Hard Prompts13771364
DTBench62.8%76.3%
LMCA18.1%22.1%
Epoch Capabilities Index133.94139.93
SimpleBench—22.1%
CritPt—1.1%
Chess Puzzles—20%
Mystery Game Puzzles—2%
Surface Evolver Bench—25%

Math gpt-oss-120b leads

Gemini 2.5 Flash-Lite: 38.0 (#144), gpt-oss-120b: 52.5 (#50)

Math benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
Omni-MATH48%68.8%
LMArena Math13731389
OTIS Mock AIME 2024-2025—88.9%

Knowledge gpt-oss-120b leads

Gemini 2.5 Flash-Lite: 32.5 (#210), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
MMLU-Pro53.7%79.5%
Vectara Hallucination Rate3.3%14.2%
GPQA (HELM)30.9%68.4%
LMArena Expert13731356
GPQA Diamond—75.8%
Confabulations—15.7%

Multimodal Not comparable

Gemini 2.5 Flash-Lite: 29.1 (#114), gpt-oss-120b: —

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
LMArena Vision1198—
VPCT30%—

Multilingual Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 49.3 (#134), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
LMArena Non-English13691351
LMArena Chinese14041385
LMArena French13881369
LMArena German13891353
LMArena Japanese13591331
LMArena Korean13601282
LMArena Russian13731343
LMArena Spanish13961389

Instruction Following Too close to call

Gemini 2.5 Flash-Lite: 70.0 (#168), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
IFEval81%83.6%
LMArena Instruction Following13671318

Long Context Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 33.3 (#262), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
Fiction.LiveBench47.2%44.4%
LMArena Longer Query13731319

Writing & Preference Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 56.8 (#135), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-Litegpt-oss-120b
LMArena Text13791365
LMArena Creative Writing13671275
WildBench81.8%84.5%
LMArena Multi-Turn13661340
Short-Story Creative Writing—77.1%
EQ-Bench Creative Writing—961

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than gpt-oss-120b?

Gemini 2.5 Flash-Lite and gpt-oss-120b score almost the same on the Noometry Index (37.0 vs 36.3), so choose on price, context window or the category you care about most.

Which is cheaper, Gemini 2.5 Flash-Lite or gpt-oss-120b?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Gemini 2.5 Flash-Lite lists at $0.10 and $0.40.

Is Gemini 2.5 Flash-Lite or gpt-oss-120b better for coding?

Gemini 2.5 Flash-Lite scores higher on coding benchmarks: 38.5 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 131K.

How many benchmarks do Gemini 2.5 Flash-Lite and gpt-oss-120b share?

30 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper