Model comparison

Llama 2-13B vs Llama 3.2 3B

Llama 2-13B and Llama 3.2 3B score almost the same on the Noometry Index (29.6 vs 28.9), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Llama 2-13B Meta

29.6

Rank #309 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 2-13B scores higher in 3 categories and Llama 3.2 3B in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.2 3B leads 21.0 to 12.8.

Side by side

Llama 2-13B and Llama 3.2 3B specifications
Llama 2-13BLlama 3.2 3B
ProviderMetaMeta
Noometry Index29.628.9
Released2023-07-182024-09-24
WeightsOpenOpen
Context window—131K
Max output—118K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.33
Results tracked3218

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 2-13B leads

Llama 2-13B: 30.9 (#291), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Coding10621098
BigCodeBench Instruct—23.4%
BigCodeBench Complete—28.3%

Agentic & Tool Use Not comparable

Llama 2-13B: —, Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
Berkeley Function Calling Leaderboard—21.9%
BALROG—10.1%

Reasoning Llama 3.2 3B leads

Llama 2-13B: 12.8 (#337), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Hard Prompts10511095
Chess Puzzles0%—
DTBench42.2%—
BIG-Bench Hard58.2%—
Epoch Capabilities Index106.17—
HellaSwag80.7%—
LAMBADA76.5%—
PIQA80.8%—
WinoGrande72.8%—

Math Llama 3.2 3B leads

Llama 2-13B: 31.1 (#229), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Math10651126
GSM8K36.9%—

Knowledge Llama 3.2 3B leads

Llama 2-13B: 28.1 (#249), Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Expert10301090
ARC (AI2) Challenge60.3%—
BoolQ82.4%—
MMLU55.6%—
OpenBookQA57%—
TriviaQA79.6%—

Multimodal Not comparable

Llama 2-13B: —, Llama 3.2 3B: —

Multimodal benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
ScienceQA55.8%—

Multilingual Too close to call

Llama 2-13B: 26.5 (#279), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Non-English10241019
LMArena Chinese10011017
LMArena German10091056
LMArena Russian1055949
LMArena French1044—
LMArena Japanese894—
LMArena Korean953—
LMArena Spanish1087—

Instruction Following Llama 3.2 3B leads

Llama 2-13B: 53.3 (#287), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Instruction Following10451089

Long Context Llama 3.2 3B leads

Llama 2-13B: 32.3 (#269), Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Longer Query10641100

Writing & Preference Llama 2-13B leads

Llama 2-13B: 29.8 (#289), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkLlama 2-13BLlama 3.2 3B
LMArena Text10841110
LMArena Creative Writing10471094
LMArena Multi-Turn10501105
EQ-Bench Creative Writing—595

Frequently asked questions

Is Llama 2-13B better than Llama 3.2 3B?

Llama 2-13B and Llama 3.2 3B score almost the same on the Noometry Index (29.6 vs 28.9), so choose on price, context window or the category you care about most.

Is Llama 2-13B or Llama 3.2 3B better for coding?

Llama 2-13B scores higher on coding benchmarks: 30.9 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 2-13B and Llama 3.2 3B share?

13 benchmarks have published results for both models. Llama 2-13B has 32 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper