Model comparison

Gemini 2.0 Flash (Feb 2025) vs o3-mini

o3-mini is the stronger model overall, scoring 36.7 to 35.1 on the Noometry Index.

Last verified . 38 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

o3-mini OpenAI

36.7

Rank #212 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 3 categories and o3-mini in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where o3-mini leads 40.8 to 28.4.
  • The biggest single-benchmark swing is CadEval: 30% for Gemini 2.0 Flash (Feb 2025) and 54% for o3-mini.

Side by side

Gemini 2.0 Flash (Feb 2025) and o3-mini specifications
Gemini 2.0 Flash (Feb 2025)o3-mini
ProviderGoogleOpenAI
Noometry Index35.136.7
Released2024-12-062024-12-20
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$1.10
Output $ / M tokens—$4.40
Results tracked5451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-mini leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), o3-mini: 40.8 (#132)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
Aider Polyglot38.2%60.4%
WeirdML25.8%43.7%
LiveBench Coding63.4%82.7%
LMArena Coding13501378
CadEval30%54%
SWE-bench Verified (bash only)13.5%—
SciCode—39.8%
GSO—1.3%
BigCodeBench Instruct45.9%—
BigCodeBench Complete59.9%—

Agentic & Tool Use o3-mini leads

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), o3-mini: 29.6 (#84)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
TheAgentCompany11.4%—
Cybench—22.5%

Reasoning o3-mini leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), o3-mini: 16.3 (#305)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
ARC-AGI-21.3%3%
SimpleBench31.1%22.8%
LiveBench Reasoning78.2%89.6%
LMArena Hard Prompts13461366
DTBench63.2%68.8%
LiveBench Data Analysis69.4%70.6%
Epoch Capabilities Index135.36140.34
LiveBench66.9%75.9%
Kagi LLM Benchmark37.8%—
ARC-AGI-1—34.5%
CritPt—0.3%
Chess Puzzles—17%
EnigmaEval1.1%—
Mystery Game Puzzles—7%
LMCA—19%
ForecastBench—59.6

Math Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), o3-mini: 28.1 (#244)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
OTIS Mock AIME 2024-202557.8%76.9%
LiveBench Math75.8%77.3%
LMArena Math13521396
MATH Level 582.2%96.5%
FrontierMath (Feb 2025 set)1.7%12.4%
FrontierMath (Tiers 1-3)—18.6%
FrontierMath Tier 4—0%
Omni-MATH45.9%—
FrontierMath Tier 4 (v1)—4.2%

Knowledge o3-mini leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), o3-mini: 38.3 (#146)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
GPQA Diamond64.1%77%
Confabulations12.4%17.9%
LMArena Expert13391364
Humanity's Last Exam6.6%—
SimpleQA Verified—15.3%
MMLU-Pro73.7%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), o3-mini: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
LMArena Vision1158—
GeoBench77%—

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), o3-mini: 45.7 (#164)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
LMArena Non-English13421319
LMArena Chinese13731379
LMArena French13911334
LMArena German13531303
LMArena Japanese12941286
LMArena Korean13131314
LMArena Russian13511304
LMArena Spanish13631321

Instruction Following Too close to call

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), o3-mini: 75.1 (#72)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
LiveBench Instruction Following85.8%84.4%
LMArena Instruction Following13361337
IFEval84.1%—

Long Context Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), o3-mini: 33.8 (#256)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
Fiction.LiveBench61.1%50%
LMArena Longer Query13441343

Writing & Preference Too close to call

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), o3-mini: 50.3 (#182)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)o3-mini
LMArena Text13541337
LMArena Creative Writing13401286
Short-Story Creative Writing73.8%61.7%
LMArena Multi-Turn13501320
LiveBench Language51.3%50.7%
EQ-Bench Creative Writing1128—
WildBench80%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than o3-mini?

o3-mini is the stronger model overall, scoring 36.7 to 35.1 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or o3-mini better for coding?

o3-mini scores higher on coding benchmarks: 40.8 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and o3-mini share?

38 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and o3-mini has 51.

Related comparisons

Go deeper