Model comparison

Llama 13b vs Llama 3.1-405B

Llama 3.1-405B is the stronger model overall, scoring 30.7 to 24.4 on the Noometry Index.

Last verified . 16 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Llama 13b scores higher in 1 category and Llama 3.1-405B in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.1-405B leads 65.9 to 36.7.

Side by side

Llama 13b and Llama 3.1-405B specifications
Llama 13bLlama 3.1-405B
ProviderMetaMeta
Noometry Index24.430.7
Released2023-02-242024-07-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2142

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Llama 13b: 21.4 (#337), Llama 3.1-405B: 33.1 (#262)

Coding benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Coding6831291
WeirdML—21.4%

Agentic & Tool Use Not comparable

Llama 13b: —, Llama 3.1-405B: 21.0 (#140)

Agentic & Tool Use benchmarks
BenchmarkLlama 13bLlama 3.1-405B
TheAgentCompany—7.4%
Cybench—7.5%

Reasoning Llama 3.1-405B leads

Llama 13b: 14.0 (#329), Llama 3.1-405B: 16.8 (#300)

Reasoning benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Hard Prompts7281269
BIG-Bench Hard37.9%82.9%
Epoch Capabilities Index100.58128.75
HellaSwag79.2%89.2%
PIQA80.1%85.9%
WinoGrande73%89.2%
SimpleBench—23%
Kagi LLM Benchmark—45%
DTBench—61.4%
ForecastBench—59.9
LAMBADA75.2%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Llama 3.1-405B: 18.4 (#290)

Math benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Math8381281
OTIS Mock AIME 2024-2025—9.7%
Omni-MATH—24.9%
MATH Level 5—49.8%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Llama 3.1-405B: 30.4 (#227)

Knowledge benchmarks
BenchmarkLlama 13bLlama 3.1-405B
ARC (AI2) Challenge52.7%95.3%
MMLU47.7%84.5%
TriviaQA77.9%82.7%
GPQA Diamond—50.9%
MMLU-Pro—72.3%
Confabulations—17.6%
GPQA (HELM)—52.2%
LMArena Expert—1243
BoolQ78.7%—
OpenBookQA56.4%—

Multimodal Not comparable

Llama 13b: —, Llama 3.1-405B: —

Multimodal benchmarks
BenchmarkLlama 13bLlama 3.1-405B
ScienceQA43.3%—

Multilingual Llama 3.1-405B leads

Llama 13b: 16.6 (#297), Llama 3.1-405B: 40.7 (#214)

Multilingual benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Non-English8191248
LMArena Chinese—1242
LMArena French—1279
LMArena German—1252
LMArena Japanese—1208
LMArena Korean—1184
LMArena Russian—1265
LMArena Spanish—1260

Instruction Following Llama 3.1-405B leads

Llama 13b: 36.7 (#305), Llama 3.1-405B: 65.9 (#214)

Instruction Following benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Instruction Following7811259
IFEval—81.1%

Long Context Not comparable

Llama 13b: —, Llama 3.1-405B: 38.4 (#197)

Long Context benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Longer Query—1266

Writing & Preference Llama 3.1-405B leads

Llama 13b: 13.8 (#312), Llama 3.1-405B: 38.9 (#251)

Writing & Preference benchmarks
BenchmarkLlama 13bLlama 3.1-405B
LMArena Text8341284
LMArena Creative Writing7941262
LMArena Multi-Turn7531297
EQ-Bench Creative Writing—870
WildBench—78.3%

Frequently asked questions

Is Llama 13b better than Llama 3.1-405B?

Llama 3.1-405B is the stronger model overall, scoring 30.7 to 24.4 on the Noometry Index.

Is Llama 13b or Llama 3.1-405B better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Llama 3.1-405B share?

16 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Llama 3.1-405B has 42.

Related comparisons

Go deeper