Model comparison

Llama 13b vs Llama 3.2 3B

Llama 3.2 3B is the stronger model overall, scoring 28.9 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 0 categories and Llama 3.2 3B in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.2 3B leads 56.0 to 36.7.

Side by side

Llama 13b and Llama 3.2 3B specifications
Llama 13bLlama 3.2 3B
ProviderMetaMeta
Noometry Index24.428.9
Released2023-02-242024-09-24
WeightsOpenOpen
Context window—131K
Max output—118K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.33
Results tracked2118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.2 3B leads

Llama 13b: 21.4 (#337), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Coding6831098
BigCodeBench Instruct—23.4%
BigCodeBench Complete—28.3%

Agentic & Tool Use Not comparable

Llama 13b: —, Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkLlama 13bLlama 3.2 3B
Berkeley Function Calling Leaderboard—21.9%
BALROG—10.1%

Reasoning Llama 3.2 3B leads

Llama 13b: 14.0 (#329), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Hard Prompts7281095
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 3.2 3B leads

Llama 13b: 26.7 (#256), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Math8381126
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Expert—1090
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Llama 3.2 3B: —

Multimodal benchmarks
BenchmarkLlama 13bLlama 3.2 3B
ScienceQA43.3%—

Multilingual Llama 3.2 3B leads

Llama 13b: 16.6 (#297), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Non-English8191019
LMArena Chinese—1017
LMArena German—1056
LMArena Russian—949

Instruction Following Llama 3.2 3B leads

Llama 13b: 36.7 (#305), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Instruction Following7811089

Long Context Not comparable

Llama 13b: —, Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Longer Query—1100

Writing & Preference Llama 3.2 3B leads

Llama 13b: 13.8 (#312), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkLlama 13bLlama 3.2 3B
LMArena Text8341110
LMArena Creative Writing7941094
LMArena Multi-Turn7531105
EQ-Bench Creative Writing—595

Frequently asked questions

Is Llama 13b better than Llama 3.2 3B?

Llama 3.2 3B is the stronger model overall, scoring 28.9 to 24.4 on the Noometry Index.

Is Llama 13b or Llama 3.2 3B better for coding?

Llama 3.2 3B scores higher on coding benchmarks: 27.6 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Llama 3.2 3B share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper