Model comparison

Gemini 1.5 Flash 8B vs o1

o1 is the stronger model overall, scoring 40.9 to 29.9 on the Noometry Index.

Last verified . 21 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 0 categories and o1 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o1 leads 41.5 to 16.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.6% for Gemini 1.5 Flash 8B and 73.3% for o1.

Side by side

Gemini 1.5 Flash 8B and o1 specifications
Gemini 1.5 Flash 8Bo1
ProviderGoogleOpenAI
Noometry Index29.940.9
Released2024-10-032024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked2152

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Gemini 1.5 Flash 8B: 35.5 (#225), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Coding12181367
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Gemini 1.5 Flash 8B: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Gemini 1.5 Flash 8B: 20.0 (#244), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Hard Prompts12091371
DTBench50%74.7%
SimpleBench—41.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math o1 leads

Gemini 1.5 Flash 8B: 14.2 (#302), o1: 36.1 (#175)

Math benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
OTIS Mock AIME 2024-20254.6%73.3%
LMArena Math12071388
FrontierMath (Tiers 1-3)—14.7%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Gemini 1.5 Flash 8B: 16.0 (#289), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
GPQA Diamond33%76.8%
LMArena Expert11851361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal o1 leads

Gemini 1.5 Flash 8B: 28.2 (#115), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Vision10441168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Gemini 1.5 Flash 8B: 38.5 (#229), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Non-English12151358
LMArena Chinese12311394
LMArena French12341344
LMArena German12061337
LMArena Japanese11501346
LMArena Korean11401396
LMArena Russian12361356
LMArena Spanish12121345

Instruction Following o1 leads

Gemini 1.5 Flash 8B: 62.8 (#236), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Instruction Following11991367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Gemini 1.5 Flash 8B: 37.0 (#225), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Longer Query12191378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Gemini 1.5 Flash 8B: 42.8 (#232), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8Bo1
LMArena Text12261366
LMArena Creative Writing12181348
LMArena Multi-Turn11851369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Gemini 1.5 Flash 8B better than o1?

o1 is the stronger model overall, scoring 40.9 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 35.5 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and o1 share?

21 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper