Model comparison

Gemini 1.0 Pro vs Llama 3-8B

Gemini 1.0 Pro is the stronger model overall, scoring 27.3 to 25.5 on the Noometry Index.

Last verified . 24 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 6 categories and Llama 3-8B in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemini 1.0 Pro leads 15.6 to 7.8.
  • The biggest single-benchmark swing is GPQA Diamond: 34% for Gemini 1.0 Pro and 26.1% for Llama 3-8B.
  • Llama 3-8B has downloadable open weights; the other is API-only.

Side by side

Gemini 1.0 Pro and Llama 3-8B specifications
Gemini 1.0 ProLlama 3-8B
ProviderGoogleMeta
Noometry Index27.325.5
Released2023-12-132024-04-18
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2434

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.0 Pro leads

Gemini 1.0 Pro: 32.2 (#275), Llama 3-8B: 31.0 (#289)

Coding benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
LMArena Coding11081152
HumanEval+55.5%56.7%
MBPP+61.4%54.8%
BigCodeBench Instruct—31.9%
BigCodeBench Complete—36.9%

Reasoning Gemini 1.0 Pro leads

Gemini 1.0 Pro: 17.1 (#296), Llama 3-8B: 14.3 (#326)

Reasoning benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
LMArena Hard Prompts11091133
DTBench45.9%43.9%
Epoch Capabilities Index117.04116.45
Chess Puzzles—0%
Adversarial NLI—57.3%
ForecastBench—58.6
WinoGrande—75.7%

Math Too close to call

Gemini 1.0 Pro: 9.3 (#321), Llama 3-8B: 8.8 (#323)

Math benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
OTIS Mock AIME 2024-20251.1%1.9%
LMArena Math11321151
MATH Level 511.2%6.1%

Knowledge Gemini 1.0 Pro leads

Gemini 1.0 Pro: 15.6 (#291), Llama 3-8B: 7.8 (#308)

Knowledge benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
GPQA Diamond34%26.1%
LMArena Expert10591113
MMLU70%68.8%
ARC (AI2) Challenge—82.8%
OpenBookQA—82.6%
TriviaQA—67.7%

Multilingual Gemini 1.0 Pro leads

Gemini 1.0 Pro: 33.4 (#252), Llama 3-8B: 30.8 (#261)

Multilingual benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
LMArena Non-English11381098
LMArena Chinese11241076
LMArena French11451159
LMArena German11251104
LMArena Japanese1023967
LMArena Russian11861109
LMArena Spanish11191173
LMArena Korean—1004

Instruction Following Too close to call

Gemini 1.0 Pro: 57.6 (#267), Llama 3-8B: 58.4 (#260)

Instruction Following benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
LMArena Instruction Following11141127

Long Context Too close to call

Gemini 1.0 Pro: 34.3 (#249), Llama 3-8B: 34.2 (#251)

Long Context benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
LMArena Longer Query11321128

Writing & Preference Llama 3-8B leads

Gemini 1.0 Pro: 36.0 (#264), Llama 3-8B: 37.5 (#256)

Writing & Preference benchmarks
BenchmarkGemini 1.0 ProLlama 3-8B
LMArena Text11491166
LMArena Creative Writing11311150
LMArena Multi-Turn11391152

Frequently asked questions

Is Gemini 1.0 Pro better than Llama 3-8B?

Gemini 1.0 Pro is the stronger model overall, scoring 27.3 to 25.5 on the Noometry Index.

Is Gemini 1.0 Pro or Llama 3-8B better for coding?

Gemini 1.0 Pro scores higher on coding benchmarks: 32.2 versus 31.0 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and Llama 3-8B share?

24 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and Llama 3-8B has 34.

Related comparisons

Go deeper