Model comparison

Gemini 1.5 Flash 8B vs GPT-4

Gemini 1.5 Flash 8B and GPT-4 score almost the same on the Noometry Index (29.9 vs 29.1), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 4 categories and GPT-4 in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash 8B leads 42.8 to 34.9.
  • The biggest single-benchmark swing is DTBench: 50% for Gemini 1.5 Flash 8B and 62.7% for GPT-4.

Side by side

Gemini 1.5 Flash 8B and GPT-4 specifications
Gemini 1.5 Flash 8BGPT-4
ProviderGoogleOpenAI
Noometry Index29.929.1
Released2024-10-032023-03-14
WeightsProprietaryProprietary
Context window—8K
Max output—8K
Input $ / M tokens—$30
Output $ / M tokens—$60
Results tracked2138

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 35.5 (#225), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Coding12181254
WeirdML—12.4%
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
HumanEval+—79.3%

Agentic & Tool Use Not comparable

Gemini 1.5 Flash 8B: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
METR Time Horizons—36.1%

Reasoning Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 20.0 (#244), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Hard Prompts12091241
DTBench50%62.7%
Chess Puzzles—4%
Mystery Game Puzzles—12%
LMCA—17.1%
BIG-Bench Hard—75.1%
Epoch Capabilities Index—125.89
ForecastBench—57.8
HellaSwag—95.3%
WinoGrande—87.5%

Math Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 14.2 (#302), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
OTIS Mock AIME 2024-20254.6%1.1%
LMArena Math12071269
MATH Level 5—23%
GSM8K—92%

Knowledge GPT-4 leads

Gemini 1.5 Flash 8B: 16.0 (#289), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
GPQA Diamond33%35.7%
LMArena Expert11851211
MMLU—86.4%
TriviaQA—84.8%

Multimodal Not comparable

Gemini 1.5 Flash 8B: 28.2 (#115), GPT-4: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Vision1044—

Multilingual GPT-4 leads

Gemini 1.5 Flash 8B: 38.5 (#229), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Non-English12151246
LMArena Chinese12311242
LMArena French12341283
LMArena German12061251
LMArena Japanese11501209
LMArena Korean11401184
LMArena Russian12361251
LMArena Spanish12121261

Instruction Following GPT-4 leads

Gemini 1.5 Flash 8B: 62.8 (#236), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Instruction Following11991241

Long Context Too close to call

Gemini 1.5 Flash 8B: 37.0 (#225), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Longer Query12191244

Writing & Preference Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 42.8 (#232), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BGPT-4
LMArena Text12261263
LMArena Creative Writing12181244
LMArena Multi-Turn11851257
EQ-Bench Creative Writing—752

Frequently asked questions

Is Gemini 1.5 Flash 8B better than GPT-4?

Gemini 1.5 Flash 8B and GPT-4 score almost the same on the Noometry Index (29.9 vs 29.1), so choose on price, context window or the category you care about most.

Is Gemini 1.5 Flash 8B or GPT-4 better for coding?

Gemini 1.5 Flash 8B scores higher on coding benchmarks: 35.5 versus 31.6 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and GPT-4 share?

20 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper