Model comparison

Gemini 1.5 Flash 8B vs Llama 2-13B

Gemini 1.5 Flash 8B and Llama 2-13B score almost the same on the Noometry Index (29.9 vs 29.6), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

Llama 2-13B Meta

29.6

Rank #309 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 6 categories and Llama 2-13B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 2-13B leads 31.1 to 14.2.
  • The biggest single-benchmark swing is DTBench: 50% for Gemini 1.5 Flash 8B and 42.2% for Llama 2-13B.
  • Llama 2-13B has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash 8B and Llama 2-13B specifications
Gemini 1.5 Flash 8BLlama 2-13B
ProviderGoogleMeta
Noometry Index29.929.6
Released2024-10-032023-07-18
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2132

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 35.5 (#225), Llama 2-13B: 30.9 (#291)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Coding12181062

Reasoning Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 20.0 (#244), Llama 2-13B: 12.8 (#337)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Hard Prompts12091051
DTBench50%42.2%
Chess Puzzles—0%
BIG-Bench Hard—58.2%
Epoch Capabilities Index—106.17
HellaSwag—80.7%
LAMBADA—76.5%
PIQA—80.8%
WinoGrande—72.8%

Math Llama 2-13B leads

Gemini 1.5 Flash 8B: 14.2 (#302), Llama 2-13B: 31.1 (#229)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Math12071065
OTIS Mock AIME 2024-20254.6%—
GSM8K—36.9%

Knowledge Llama 2-13B leads

Gemini 1.5 Flash 8B: 16.0 (#289), Llama 2-13B: 28.1 (#249)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Expert11851030
GPQA Diamond33%—
ARC (AI2) Challenge—60.3%
BoolQ—82.4%
MMLU—55.6%
OpenBookQA—57%
TriviaQA—79.6%

Multimodal Not comparable

Gemini 1.5 Flash 8B: 28.2 (#115), Llama 2-13B: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Vision1044—
ScienceQA—55.8%

Multilingual Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 38.5 (#229), Llama 2-13B: 26.5 (#279)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Non-English12151024
LMArena Chinese12311001
LMArena French12341044
LMArena German12061009
LMArena Japanese1150894
LMArena Korean1140953
LMArena Russian12361055
LMArena Spanish12121087

Instruction Following Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 62.8 (#236), Llama 2-13B: 53.3 (#287)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Instruction Following11991045

Long Context Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 37.0 (#225), Llama 2-13B: 32.3 (#269)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Longer Query12191064

Writing & Preference Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 42.8 (#232), Llama 2-13B: 29.8 (#289)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BLlama 2-13B
LMArena Text12261084
LMArena Creative Writing12181047
LMArena Multi-Turn11851050

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Llama 2-13B?

Gemini 1.5 Flash 8B and Llama 2-13B score almost the same on the Noometry Index (29.9 vs 29.6), so choose on price, context window or the category you care about most.

Is Gemini 1.5 Flash 8B or Llama 2-13B better for coding?

Gemini 1.5 Flash 8B scores higher on coding benchmarks: 35.5 versus 30.9 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and Llama 2-13B share?

18 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Llama 2-13B has 32.

Related comparisons

Go deeper