Model comparison

Gemini 2.0 Flash (Feb 2025) vs gpt-oss-20b

Gemini 2.0 Flash (Feb 2025) is the stronger model overall, scoring 35.1 to 32.5 on the Noometry Index.

Last verified . 28 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 5 categories and gpt-oss-20b in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Gemini 2.0 Flash (Feb 2025) leads 28.1 to 9.3.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 37.8% for Gemini 2.0 Flash (Feb 2025) and 53.2% for gpt-oss-20b.
  • gpt-oss-20b has downloadable open weights; the other is API-only.

Side by side

Gemini 2.0 Flash (Feb 2025) and gpt-oss-20b specifications
Gemini 2.0 Flash (Feb 2025)gpt-oss-20b
ProviderGoogleOpenAI
Noometry Index35.132.5
Released2024-12-062025-08-05
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.018
Output $ / M tokens—$0.09
Results tracked5434

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding gpt-oss-20b leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), gpt-oss-20b: 37.6 (#192)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
WeirdML25.8%40.9%
LMArena Coding13501306
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
SciCode—34.4%
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—
ALE-Bench—566.05

Agentic & Tool Use Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), gpt-oss-20b: 9.3 (#154)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
Terminal-Bench—3.4%
TheAgentCompany11.4%—

Reasoning gpt-oss-20b leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), gpt-oss-20b: 19.3 (#261)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
Kagi LLM Benchmark37.8%53.2%
LMArena Hard Prompts13461274
DTBench63.2%68%
Epoch Capabilities Index135.36137.82
ARC-AGI-21.3%—
SimpleBench31.1%—
CritPt—1.4%
Chess Puzzles—4%
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
LiveBench Data Analysis69.4%—
LMCA—14.5%
LiveBench66.9%—

Math gpt-oss-20b leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), gpt-oss-20b: 39.4 (#103)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
OTIS Mock AIME 2024-202557.8%65.3%
Omni-MATH45.9%56.5%
LMArena Math13521317
LiveBench Math75.8%—
MATH Level 582.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge gpt-oss-20b leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), gpt-oss-20b: 34.6 (#195)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
GPQA Diamond64.1%60.8%
MMLU-Pro73.7%74%
GPQA (HELM)55.6%59.4%
LMArena Expert13391258
Humanity's Last Exam6.6%—
Confabulations12.4%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), gpt-oss-20b: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
LMArena Vision1158—
GeoBench77%—

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), gpt-oss-20b: 42.2 (#197)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
LMArena Non-English13421268
LMArena Chinese13731314
LMArena German13531255
LMArena Japanese12941244
LMArena Korean13131236
LMArena Russian13511278
LMArena Spanish13631267
LMArena French1391—

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), gpt-oss-20b: 61.8 (#240)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
IFEval84.1%73.2%
LMArena Instruction Following13361236
LiveBench Instruction Following85.8%—

Long Context Too close to call

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), gpt-oss-20b: 37.9 (#209)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
LMArena Longer Query13441250
Fiction.LiveBench61.1%—

Writing & Preference Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), gpt-oss-20b: 35.5 (#265)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)gpt-oss-20b
LMArena Text13541287
LMArena Creative Writing13401201
EQ-Bench Creative Writing1128666
WildBench80%73.7%
LMArena Multi-Turn13501268
Short-Story Creative Writing73.8%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than gpt-oss-20b?

Gemini 2.0 Flash (Feb 2025) is the stronger model overall, scoring 35.1 to 32.5 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or gpt-oss-20b better for coding?

gpt-oss-20b scores higher on coding benchmarks: 37.6 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and gpt-oss-20b share?

28 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and gpt-oss-20b has 34.

Related comparisons

Go deeper