Model comparison

Gemini 1.5 Flash (May 2024) vs Llama 13b

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 24.4 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 5 categories and Llama 13b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash (May 2024) leads 48.7 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Llama 13b specifications
Gemini 1.5 Flash (May 2024)Llama 13b
ProviderGoogleMeta
Noometry Index33.224.4
Released2024-05-142023-02-24
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4221

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Coding1261683
WeirdML24.9%—
BigCodeBench Instruct43.5%—
BigCodeBench Complete55.1%—
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Not comparable

Gemini 1.5 Flash (May 2024): 26.6 (#102), Llama 13b: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
BALROG14.6%—

Reasoning Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Hard Prompts1257728
Epoch Capabilities Index129.36100.58
PIQA87.5%80.1%
DTBench53.8%—
BIG-Bench Hard—37.9%
ForecastBench53.9—
HellaSwag—79.2%
LAMBADA—75.2%
WinoGrande—73%

Math Llama 13b leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Math1269838
GSM8K82.4%20.6%
OTIS Mock AIME 2024-202516.3%—
Omni-MATH30.4%—
MATH Level 561.9%—
FrontierMath (Feb 2025 set)0%—

Knowledge Not comparable

Gemini 1.5 Flash (May 2024): 26.2 (#260), Llama 13b: —

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
BoolQ85.8%78.7%
MMLU77.9%47.7%
GPQA Diamond47.3%—
MMLU-Pro67.8%—
GPQA (HELM)43.7%—
LMArena Expert1233—
ARC (AI2) Challenge—52.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), Llama 13b: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—
ScienceQA—43.3%

Multilingual Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Non-English1278819
LMArena Chinese1295—
LMArena French1258—
LMArena German1262—
LMArena Japanese1252—
LMArena Korean1221—
LMArena Russian1288—
LMArena Spanish1243—

Instruction Following Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Instruction Following1258781
IFEval83.1%—

Long Context Not comparable

Gemini 1.5 Flash (May 2024): 39.0 (#187), Llama 13b: —

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Longer Query1284—

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Llama 13b
LMArena Text1287834
LMArena Creative Writing1285794
LMArena Multi-Turn1253753
WildBench79.2%—

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Llama 13b?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 24.4 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Llama 13b better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 21.4 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Llama 13b share?

13 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper