Model comparison

Llama 3.2 3B vs Phi 3 Small 8k Instruct

Llama 3.2 3B and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (28.9 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.2 3B scores higher in 5 categories and Phi 3 Small 8k Instruct in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Phi 3 Small 8k Instruct leads 31.1 to 24.7.

Side by side

Llama 3.2 3B and Phi 3 Small 8k Instruct specifications
Llama 3.2 3BPhi 3 Small 8k Instruct
ProviderMetaMicrosoft
Noometry Index28.929.3
Released2024-09-242024-04-23
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1832

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.2 3B: 27.6 (#319), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Coding10981101
BigCodeBench Instruct23.4%—
LiveBench Coding—20.3%
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Phi 3 Small 8k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Hard Prompts10951100
LiveBench Reasoning—15.9%
LiveBench Data Analysis—30.3%
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
HellaSwag—77%
LiveBench—24%
WinoGrande—81.5%

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Math11261151
LiveBench Math—17.6%

Knowledge Too close to call

Llama 3.2 3B: 29.7 (#235), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Expert10901067
ARC (AI2) Challenge—90.7%
MMLU—75.7%
OpenBookQA—88%
TriviaQA—58.1%

Multilingual Phi 3 Small 8k Instruct leads

Llama 3.2 3B: 26.2 (#281), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Non-English10191058
LMArena Chinese10171061
LMArena German10561080
LMArena Russian9491111
LMArena French—1135
LMArena Japanese—966
LMArena Korean—894
LMArena Spanish—1111

Instruction Following Llama 3.2 3B leads

Llama 3.2 3B: 56.0 (#275), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Instruction Following10891087
LiveBench Instruction Following—47.2%

Long Context Too close to call

Llama 3.2 3B: 33.4 (#261), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Longer Query11001088

Writing & Preference Phi 3 Small 8k Instruct leads

Llama 3.2 3B: 24.7 (#307), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BPhi 3 Small 8k Instruct
LMArena Text11101110
LMArena Creative Writing10941083
LMArena Multi-Turn11051068
EQ-Bench Creative Writing595—
LiveBench Language—12.9%

Frequently asked questions

Is Llama 3.2 3B better than Phi 3 Small 8k Instruct?

Llama 3.2 3B and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (28.9 vs 29.3), so choose on price, context window or the category you care about most.

Is Llama 3.2 3B or Phi 3 Small 8k Instruct better for coding?

They score almost the same on coding (27.6 vs 27.9); test both on your own repository before choosing.

How many benchmarks do Llama 3.2 3B and Phi 3 Small 8k Instruct share?

13 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper