Model comparison

Llama 3.2 3B vs phi-3-medium 14B

Llama 3.2 3B and phi-3-medium 14B score almost the same on the Noometry Index (28.9 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. Llama 3.2 3B scores higher in 2 categories and phi-3-medium 14B in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3.2 3B leads 29.7 to 9.1.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 48.7% for phi-3-medium 14B.

Side by side

Llama 3.2 3B and phi-3-medium 14B specifications
Llama 3.2 3Bphi-3-medium 14B
ProviderMetaMicrosoft
Noometry Index28.929.7
Released2024-09-242024-04-23
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1813

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

Llama 3.2 3B: 27.6 (#319), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
BigCodeBench Instruct23.4%37.6%
BigCodeBench Complete28.3%48.7%
LMArena Coding1098—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), phi-3-medium 14B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Not comparable

Llama 3.2 3B: 21.0 (#228), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
LMArena Hard Prompts1095—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
Epoch Capabilities Index—121.23
HellaSwag—82.4%
WinoGrande—81.5%

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
LMArena Math1126—
MATH Level 5—17.6%

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
GPQA Diamond—27.6%
LMArena Expert1090—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

Llama 3.2 3B: 26.2 (#281), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
LMArena Non-English1019—
LMArena Chinese1017—
LMArena German1056—
LMArena Russian949—

Instruction Following Not comparable

Llama 3.2 3B: 56.0 (#275), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
LMArena Instruction Following1089—

Long Context Not comparable

Llama 3.2 3B: 33.4 (#261), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
LMArena Longer Query1100—

Writing & Preference Not comparable

Llama 3.2 3B: 24.7 (#307), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkLlama 3.2 3Bphi-3-medium 14B
LMArena Text1110—
LMArena Creative Writing1094—
EQ-Bench Creative Writing595—
LMArena Multi-Turn1105—

Frequently asked questions

Is Llama 3.2 3B better than phi-3-medium 14B?

Llama 3.2 3B and phi-3-medium 14B score almost the same on the Noometry Index (28.9 vs 29.7), so choose on price, context window or the category you care about most.

Is Llama 3.2 3B or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and phi-3-medium 14B share?

2 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper