Model comparison

Llama 3.1-405B vs Llama 3.1-70B

Llama 3.1-405B is the stronger model overall, scoring 30.7 to 29.6 on the Noometry Index.

Last verified . 31 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Llama 3.1-405B scores higher in 7 categories and Llama 3.1-70B in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3.1-405B leads 30.4 to 24.2.
  • The biggest single-benchmark swing is MATH Level 5: 49.8% for Llama 3.1-405B and 36.7% for Llama 3.1-70B.

Side by side

Llama 3.1-405B and Llama 3.1-70B specifications
Llama 3.1-405BLlama 3.1-70B
ProviderMetaMeta
Noometry Index30.729.6
Released2024-07-232024-07-23
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.40
Output $ / M tokens—$0.40
Results tracked4235

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Llama 3.1-405B: 33.1 (#262), Llama 3.1-70B: 30.3 (#296)

Coding benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
WeirdML21.4%9%
LMArena Coding12911260
BigCodeBench Instruct—46.1%
BigCodeBench Complete—54.8%

Agentic & Tool Use Llama 3.1-70B leads

Llama 3.1-405B: 21.0 (#140), Llama 3.1-70B: 25.1 (#112)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
TheAgentCompany7.4%6.9%
Cybench7.5%—
BALROG—27.9%

Reasoning Llama 3.1-70B leads

Llama 3.1-405B: 16.8 (#300), Llama 3.1-70B: 21.6 (#220)

Reasoning benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
LMArena Hard Prompts12691241
DTBench61.4%60%
Epoch Capabilities Index128.75125.92
SimpleBench23%—
Kagi LLM Benchmark45%—
LMCA—14.8%
BIG-Bench Hard82.9%—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Llama 3.1-405B leads

Llama 3.1-405B: 18.4 (#290), Llama 3.1-70B: 13.5 (#304)

Math benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
OTIS Mock AIME 2024-20259.7%3.6%
Omni-MATH24.9%21%
LMArena Math12811252
MATH Level 549.8%36.7%

Knowledge Llama 3.1-405B leads

Llama 3.1-405B: 30.4 (#227), Llama 3.1-70B: 24.2 (#269)

Knowledge benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
GPQA Diamond50.9%44.2%
MMLU-Pro72.3%65.3%
GPQA (HELM)52.2%42.6%
LMArena Expert12431209
MMLU84.5%80.1%
Confabulations17.6%—
ARC (AI2) Challenge95.3%—
TriviaQA82.7%—

Multilingual Llama 3.1-405B leads

Llama 3.1-405B: 40.7 (#214), Llama 3.1-70B: 38.8 (#225)

Multilingual benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
LMArena Non-English12481219
LMArena Chinese12421215
LMArena French12791261
LMArena German12521222
LMArena Japanese12081132
LMArena Korean11841140
LMArena Russian12651234
LMArena Spanish12601253

Instruction Following Too close to call

Llama 3.1-405B: 65.9 (#214), Llama 3.1-70B: 65.3 (#223)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
IFEval81.1%82.1%
LMArena Instruction Following12591231

Long Context Too close to call

Llama 3.1-405B: 38.4 (#197), Llama 3.1-70B: 37.6 (#214)

Long Context benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
LMArena Longer Query12661241

Writing & Preference Llama 3.1-405B leads

Llama 3.1-405B: 38.9 (#251), Llama 3.1-70B: 35.4 (#267)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BLlama 3.1-70B
LMArena Text12841261
LMArena Creative Writing12621232
EQ-Bench Creative Writing870784
WildBench78.3%75.8%
LMArena Multi-Turn12971256

Frequently asked questions

Is Llama 3.1-405B better than Llama 3.1-70B?

Llama 3.1-405B is the stronger model overall, scoring 30.7 to 29.6 on the Noometry Index.

Is Llama 3.1-405B or Llama 3.1-70B better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 30.3 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Llama 3.1-70B share?

31 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Llama 3.1-70B has 35.

Related comparisons

Go deeper