Model comparison

Llama 2-70B vs Llama 2-7B

Llama 2-7B is the stronger model overall, scoring 29.1 to 24.4 on the Noometry Index.

Last verified . 27 shared benchmarks.

Llama 2-70B Meta

24.4

Rank #349 Confirmed

Llama 2-7B Meta

29.1

Rank #317 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Llama 2-70B scores higher in 5 categories and Llama 2-7B in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 2-7B leads 30.7 to 8.1.

Side by side

Llama 2-70B and Llama 2-7B specifications
Llama 2-70BLlama 2-7B
ProviderMetaMeta
Noometry Index24.429.1
Released2023-07-182023-07-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3529

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 2-70B leads

Llama 2-70B: 31.4 (#286), Llama 2-7B: 29.2 (#307)

Coding benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Coding10791002

Reasoning Llama 2-7B leads

Llama 2-70B: 14.4 (#325), Llama 2-7B: 15.7 (#312)

Reasoning benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Hard Prompts10731009
BIG-Bench Hard64.9%39.2%
Epoch Capabilities Index113.7999.06
HellaSwag85.3%77.2%
LAMBADA78.9%73.3%
PIQA82.8%78.8%
WinoGrande80.2%69.2%
Chess Puzzles—0%
DTBench41.6%—
CommonsenseQA 2.050%—
ForecastBench51.4—

Math Llama 2-7B leads

Llama 2-70B: 8.1 (#326), Llama 2-7B: 30.7 (#233)

Math benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Math10911042
GSM8K69.6%16.7%
OTIS Mock AIME 2024-20250%—
MATH Level 53.3%—

Knowledge Llama 2-7B leads

Llama 2-70B: 7.4 (#310), Llama 2-7B: 28.2 (#248)

Knowledge benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Expert10391036
ARC (AI2) Challenge78.3%45.9%
BoolQ88.6%77.9%
MMLU69.9%45.8%
OpenBookQA60.2%58.6%
TriviaQA87.6%73.7%
GPQA Diamond26.3%—

Multimodal Not comparable

Llama 2-70B: —, Llama 2-7B: —

Multimodal benchmarks
BenchmarkLlama 2-70BLlama 2-7B
ScienceQA—43.1%

Multilingual Llama 2-70B leads

Llama 2-70B: 27.7 (#274), Llama 2-7B: 23.8 (#293)

Multilingual benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Non-English1045973
LMArena Chinese995973
LMArena French1090970
LMArena German1041978
LMArena Russian1083995
LMArena Spanish11431007
LMArena Japanese927—
LMArena Korean964—

Instruction Following Llama 2-70B leads

Llama 2-70B: 54.9 (#278), Llama 2-7B: 50.8 (#298)

Instruction Following benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Instruction Following10711006

Long Context Llama 2-70B leads

Llama 2-70B: 32.3 (#270), Llama 2-7B: 30.4 (#287)

Long Context benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Longer Query1062999

Writing & Preference Llama 2-70B leads

Llama 2-70B: 32.3 (#279), Llama 2-7B: 28.0 (#298)

Writing & Preference benchmarks
BenchmarkLlama 2-70BLlama 2-7B
LMArena Text11151053
LMArena Creative Writing10751033
LMArena Multi-Turn10881029

Frequently asked questions

Is Llama 2-70B better than Llama 2-7B?

Llama 2-7B is the stronger model overall, scoring 29.1 to 24.4 on the Noometry Index.

Is Llama 2-70B or Llama 2-7B better for coding?

Llama 2-70B scores higher on coding benchmarks: 31.4 versus 29.2 in the Noometry coding category.

How many benchmarks do Llama 2-70B and Llama 2-7B share?

27 benchmarks have published results for both models. Llama 2-70B has 35 scored results on Noometry and Llama 2-7B has 29.

Related comparisons

Go deeper