Model comparison

Llama 3.2 3B vs Phi 3 Mini 128k Instruct

Llama 3.2 3B and Phi 3 Mini 128k Instruct score almost the same on the Noometry Index (28.9 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Llama 3.2 3B scores higher in 6 categories and Phi 3 Mini 128k Instruct in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.2 3B leads 56.0 to 51.8.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 40.6% for Phi 3 Mini 128k Instruct.

Side by side

Llama 3.2 3B and Phi 3 Mini 128k Instruct specifications
Llama 3.2 3BPhi 3 Mini 128k Instruct
ProviderMetaMicrosoft
Noometry Index28.929.7
Released2024-09-242024-04-23
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1819

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi 3 Mini 128k Instruct leads

Llama 3.2 3B: 27.6 (#319), Phi 3 Mini 128k Instruct: 28.8 (#312)

Coding benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
BigCodeBench Instruct23.4%29.6%
LMArena Coding10981039
BigCodeBench Complete28.3%40.6%

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Phi 3 Mini 128k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Phi 3 Mini 128k Instruct: 19.6 (#256)

Reasoning benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Hard Prompts10951028

Math Too close to call

Llama 3.2 3B: 32.4 (#214), Phi 3 Mini 128k Instruct: 31.6 (#222)

Math benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Math11261089

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), Phi 3 Mini 128k Instruct: 26.8 (#254)

Knowledge benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Expert1090984

Multilingual Llama 3.2 3B leads

Llama 3.2 3B: 26.2 (#281), Phi 3 Mini 128k Instruct: 25.2 (#285)

Multilingual benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Non-English10191000
LMArena Chinese10171016
LMArena German10561006
LMArena Russian9491004
LMArena French—1039
LMArena Japanese—899
LMArena Korean—856
LMArena Spanish—1059

Instruction Following Llama 3.2 3B leads

Llama 3.2 3B: 56.0 (#275), Phi 3 Mini 128k Instruct: 51.8 (#294)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Instruction Following10891023

Long Context Llama 3.2 3B leads

Llama 3.2 3B: 33.4 (#261), Phi 3 Mini 128k Instruct: 30.4 (#289)

Long Context benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Longer Query1100996

Writing & Preference Phi 3 Mini 128k Instruct leads

Llama 3.2 3B: 24.7 (#307), Phi 3 Mini 128k Instruct: 27.1 (#301)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BPhi 3 Mini 128k Instruct
LMArena Text11101050
LMArena Creative Writing10941024
LMArena Multi-Turn1105989
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Phi 3 Mini 128k Instruct?

Llama 3.2 3B and Phi 3 Mini 128k Instruct score almost the same on the Noometry Index (28.9 vs 29.7), so choose on price, context window or the category you care about most.

Is Llama 3.2 3B or Phi 3 Mini 128k Instruct better for coding?

Phi 3 Mini 128k Instruct scores higher on coding benchmarks: 28.8 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and Phi 3 Mini 128k Instruct share?

15 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Phi 3 Mini 128k Instruct has 19.

Related comparisons

Go deeper