Model comparison

Gemini 1.5 Flash 8B vs Olmo 2 0325 32b Instruct

Olmo 2 0325 32b Instruct is the stronger model overall, scoring 32.7 to 29.9 on the Noometry Index.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 5 categories and Olmo 2 0325 32b Instruct in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 2 0325 32b Instruct leads 26.8 to 14.2.
  • Olmo 2 0325 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash 8B and Olmo 2 0325 32b Instruct specifications
Gemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
ProviderGoogleAllen Institute for AI (Ai2)
Noometry Index29.932.7
Released2024-10-03—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2116

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 1.5 Flash 8B: 35.5 (#225), Olmo 2 0325 32b Instruct: 35.2 (#227)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Coding12181210

Reasoning Olmo 2 0325 32b Instruct leads

Gemini 1.5 Flash 8B: 20.0 (#244), Olmo 2 0325 32b Instruct: 23.6 (#175)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Hard Prompts12091208
DTBench50%—

Math Olmo 2 0325 32b Instruct leads

Gemini 1.5 Flash 8B: 14.2 (#302), Olmo 2 0325 32b Instruct: 26.8 (#255)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Math12071208
OTIS Mock AIME 2024-20254.6%—
Omni-MATH—16.1%

Knowledge Olmo 2 0325 32b Instruct leads

Gemini 1.5 Flash 8B: 16.0 (#289), Olmo 2 0325 32b Instruct: 19.5 (#279)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
GPQA Diamond33%—
MMLU-Pro—41.4%
GPQA (HELM)—28.7%
LMArena Expert1185—

Multimodal Not comparable

Gemini 1.5 Flash 8B: 28.2 (#115), Olmo 2 0325 32b Instruct: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Vision1044—

Multilingual Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 38.5 (#229), Olmo 2 0325 32b Instruct: 34.8 (#248)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Non-English12151160
LMArena Chinese12311192
LMArena Russian12361187
LMArena French1234—
LMArena German1206—
LMArena Japanese1150—
LMArena Korean1140—
LMArena Spanish1212—

Instruction Following Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 62.8 (#236), Olmo 2 0325 32b Instruct: 61.5 (#244)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Instruction Following11991186
IFEval—78%

Long Context Too close to call

Gemini 1.5 Flash 8B: 37.0 (#225), Olmo 2 0325 32b Instruct: 36.2 (#234)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Longer Query12191194

Writing & Preference Too close to call

Gemini 1.5 Flash 8B: 42.8 (#232), Olmo 2 0325 32b Instruct: 42.1 (#236)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BOlmo 2 0325 32b Instruct
LMArena Text12261218
LMArena Creative Writing12181199
LMArena Multi-Turn11851221
WildBench—73.4%

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Olmo 2 0325 32b Instruct?

Olmo 2 0325 32b Instruct is the stronger model overall, scoring 32.7 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or Olmo 2 0325 32b Instruct better for coding?

They score almost the same on coding (35.5 vs 35.2); test both on your own repository before choosing.

How many benchmarks do Gemini 1.5 Flash 8B and Olmo 2 0325 32b Instruct share?

11 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Olmo 2 0325 32b Instruct has 16.

Related comparisons

Go deeper