Model comparison

Gemini 3.1 Flash Lite vs o3

o3 is the stronger model overall, scoring 47.5 to 40.8 on the Noometry Index. Gemini 3.1 Flash Lite costs 6.2× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 33 shared benchmarks.

Gemini 3.1 Flash Lite Google

40.8

Rank #144 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Gemini 3.1 Flash Lite scores higher in 1 category and o3 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 41.9.
  • The biggest single-benchmark swing is Chess Puzzles: 25% for Gemini 3.1 Flash Lite and 38% for o3.
  • Gemini 3.1 Flash Lite is cheaper at $0.25 / $1.50 per million input/output tokens, against $2 / $8 for o3.
  • Gemini 3.1 Flash Lite accepts more context: 1.05M tokens versus 200K.

Side by side

Gemini 3.1 Flash Lite and o3 specifications
Gemini 3.1 Flash Liteo3
ProviderGoogleOpenAI
Noometry Index40.847.5
Released2026-03-032025-04-16
WeightsProprietaryProprietary
Context window1.05M200K
Max output66K100K
Input $ / M tokens$0.25$2
Output $ / M tokens$1.50$8
Results tracked3863

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Gemini 3.1 Flash Lite: 37.8 (#188), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGemini 3.1 Flash Liteo3
WeirdML52.2%52.4%
LMArena Coding14001408
ALE-Bench797.73933.55
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1256—
SciCode41.9%—
GSO—8.8%
CadEval—74%

Agentic & Tool Use o3 leads

Gemini 3.1 Flash Lite: 30.2 (#79), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Flash Liteo3
DeepResearch Bench37.3%45.2%
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Gemini 3.1 Flash Lite: 22.9 (#186), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGemini 3.1 Flash Liteo3
Kagi LLM Benchmark67.2%67.6%
CritPt1.1%1.4%
Chess Puzzles25%38%
EnigmaEval3%13.1%
LMArena Hard Prompts14071402
DTBench76.8%84.8%
LMCA35%39.7%
Epoch Capabilities Index144.47146.86
ForecastBench54.462.5
ARC-AGI-2—6.5%
SimpleBench—53.1%
NYT Connections (extended)8.2%—
ARC-AGI-1—60.8%
Thematic Generalization63.3%—
Mystery Game Puzzles—29%

Math o3 leads

Gemini 3.1 Flash Lite: 40.7 (#90), o3: 50.2 (#58)

Math benchmarks
BenchmarkGemini 3.1 Flash Liteo3
FrontierMath (Tiers 1-3)27.7%33.3%
OTIS Mock AIME 2024-202580%84.4%
LMArena Math14281426
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Gemini 3.1 Flash Lite: 41.9 (#104), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGemini 3.1 Flash Liteo3
GPQA Diamond81.8%81.8%
Humanity's Last Exam8.6%20.3%
LMArena Expert13981402
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
Vectara Hallucination Rate8.2%—
GPQA (HELM)—75.3%

Multimodal o3 leads

Gemini 3.1 Flash Lite: 39.4 (#60), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGemini 3.1 Flash Liteo3
LMArena Vision12401214
GeoBench—74%
VPCT—52%

Multilingual Too close to call

Gemini 3.1 Flash Lite: 52.3 (#86), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGemini 3.1 Flash Liteo3
LMArena Non-English14111401
LMArena Chinese14611437
LMArena French14241430
LMArena German14291420
LMArena Japanese14131403
LMArena Korean13921370
LMArena Russian14201406
LMArena Spanish14211395

Instruction Following Too close to call

Gemini 3.1 Flash Lite: 72.7 (#131), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGemini 3.1 Flash Liteo3
LMArena Instruction Following13771368
IFEval—86.9%

Long Context o3 leads

Gemini 3.1 Flash Lite: 42.5 (#122), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGemini 3.1 Flash Liteo3
LMArena Longer Query13941372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Gemini 3.1 Flash Lite: 60.9 (#94), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Flash Liteo3
LMArena Text14161410
LMArena Creative Writing14011359
LMArena Multi-Turn14171405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Gemini 3.1 Flash Lite better than o3?

o3 is the stronger model overall, scoring 47.5 to 40.8 on the Noometry Index. Gemini 3.1 Flash Lite costs 6.2× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Flash Lite or o3?

Gemini 3.1 Flash Lite is cheaper. It lists at $0.25 per million input tokens and $1.50 per million output tokens; o3 lists at $2 and $8.

Is Gemini 3.1 Flash Lite or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 37.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Flash Lite does, with 1.05M tokens against 200K.

How many benchmarks do Gemini 3.1 Flash Lite and o3 share?

33 benchmarks have published results for both models. Gemini 3.1 Flash Lite has 38 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper