Model comparison

Gemini 1.0 Pro vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 27.3 on the Noometry Index.

Last verified . 9 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 5 categories and Llama2 70b Steerlm Chat in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 9.3.
  • Llama2 70b Steerlm Chat has downloadable open weights; the other is API-only.

Side by side

Gemini 1.0 Pro and Llama2 70b Steerlm Chat specifications
Gemini 1.0 ProLlama2 70b Steerlm Chat
ProviderGoogleNVIDIA
Noometry Index27.331.8
Released2023-12-13—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked249

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.0 Pro leads

Gemini 1.0 Pro: 32.2 (#275), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Coding11081025
HumanEval+55.5%—
MBPP+61.4%—

Reasoning Llama2 70b Steerlm Chat leads

Gemini 1.0 Pro: 17.1 (#296), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Hard Prompts11091047
DTBench45.9%—
Epoch Capabilities Index117.04—

Math Llama2 70b Steerlm Chat leads

Gemini 1.0 Pro: 9.3 (#321), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Math11321072
OTIS Mock AIME 2024-20251.1%—
MATH Level 511.2%—

Knowledge Not comparable

Gemini 1.0 Pro: 15.6 (#291), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
GPQA Diamond34%—
LMArena Expert1059—
MMLU70%—

Multilingual Gemini 1.0 Pro leads

Gemini 1.0 Pro: 33.4 (#252), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Non-English11381063
LMArena Chinese1124—
LMArena French1145—
LMArena German1125—
LMArena Japanese1023—
LMArena Russian1186—
LMArena Spanish1119—

Instruction Following Gemini 1.0 Pro leads

Gemini 1.0 Pro: 57.6 (#267), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Instruction Following11141060

Long Context Gemini 1.0 Pro leads

Gemini 1.0 Pro: 34.3 (#249), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Longer Query1132998

Writing & Preference Gemini 1.0 Pro leads

Gemini 1.0 Pro: 36.0 (#264), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkGemini 1.0 ProLlama2 70b Steerlm Chat
LMArena Text11491098
LMArena Creative Writing11311091
LMArena Multi-Turn11391058

Frequently asked questions

Is Gemini 1.0 Pro better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 27.3 on the Noometry Index.

Is Gemini 1.0 Pro or Llama2 70b Steerlm Chat better for coding?

Gemini 1.0 Pro scores higher on coding benchmarks: 32.2 versus 29.9 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper