Model comparison

Llama 13b vs Llama 3.2 90B

Llama 3.2 90B is the stronger model overall, scoring 27.5 to 24.4 on the Noometry Index.

Last verified . 2 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Llama 13b scores higher in 1 category and Llama 3.2 90B in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 13b leads 26.7 to 11.1.

Side by side

Llama 13b and Llama 3.2 90B specifications
Llama 13bLlama 3.2 90B
ProviderMetaMeta
Noometry Index24.427.5
Released2023-02-242024-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 13b: 21.4 (#337), Llama 3.2 90B: —

Coding benchmarks
BenchmarkLlama 13bLlama 3.2 90B
LMArena Coding683—

Agentic & Tool Use Not comparable

Llama 13b: —, Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkLlama 13bLlama 3.2 90B
BALROG—27.3%

Reasoning Llama 3.2 90B leads

Llama 13b: 14.0 (#329), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkLlama 13bLlama 3.2 90B
Epoch Capabilities Index100.58125.5
EnigmaEval—0.4%
LMArena Hard Prompts728—
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkLlama 13bLlama 3.2 90B
OTIS Mock AIME 2024-2025—2.6%
LMArena Math838—
MATH Level 5—39.4%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkLlama 13bLlama 3.2 90B
MMLU47.7%80.3%
GPQA Diamond—41%
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkLlama 13bLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%
ScienceQA43.3%—

Multilingual Not comparable

Llama 13b: 16.6 (#297), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkLlama 13bLlama 3.2 90B
LMArena Non-English819—

Instruction Following Not comparable

Llama 13b: 36.7 (#305), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkLlama 13bLlama 3.2 90B
LMArena Instruction Following781—

Writing & Preference Not comparable

Llama 13b: 13.8 (#312), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkLlama 13bLlama 3.2 90B
LMArena Text834—
LMArena Creative Writing794—
LMArena Multi-Turn753—

Frequently asked questions

Is Llama 13b better than Llama 3.2 90B?

Llama 3.2 90B is the stronger model overall, scoring 27.5 to 24.4 on the Noometry Index.

How many benchmarks do Llama 13b and Llama 3.2 90B share?

2 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper