Model comparison

Llama-3.3-70B-Instruct vs Phi 3 Mini 4k Instruct June 2024

Llama-3.3-70B-Instruct and Phi 3 Mini 4k Instruct June 2024 score almost the same on the Noometry Index (30.6 vs 31.3), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Llama-3.3-70B-Instruct scores higher in 4 categories and Phi 3 Mini 4k Instruct June 2024 in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama-3.3-70B-Instruct leads 47.6 to 29.6.

Side by side

Llama-3.3-70B-Instruct and Phi 3 Mini 4k Instruct June 2024 specifications
Llama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
ProviderMetaMicrosoft
Noometry Index30.631.3
Released2024-12-06—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.32—
Results tracked4315

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama-3.3-70B-Instruct: 31.0 (#290), Phi 3 Mini 4k Instruct June 2024: 31.8 (#279)

Coding benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Coding12681093
SciCode26%—
WeirdML14.4%—
BigCodeBench Instruct46.9%—
LiveBench Coding36.6%—
BigCodeBench Complete57.5%—

Agentic & Tool Use Not comparable

Llama-3.3-70B-Instruct: 25.8 (#105), Phi 3 Mini 4k Instruct June 2024: —

Agentic & Tool Use benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
Berkeley Function Calling Leaderboard31.9%—
BALROG23%—

Reasoning Phi 3 Mini 4k Instruct June 2024 leads

Llama-3.3-70B-Instruct: 14.1 (#327), Phi 3 Mini 4k Instruct June 2024: 20.8 (#231)

Reasoning benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Hard Prompts12571087
SimpleBench19.9%—
CritPt0%—
LiveBench Reasoning50.8%—
DTBench59.5%—
LiveBench Data Analysis49.5%—
LMCA17.5%—
Epoch Capabilities Index127.33—
ForecastBench58.6—
LiveBench50.2%—

Math Phi 3 Mini 4k Instruct June 2024 leads

Llama-3.3-70B-Instruct: 15.3 (#298), Phi 3 Mini 4k Instruct June 2024: 33.0 (#208)

Math benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Math12671152
OTIS Mock AIME 2024-20255.1%—
LiveBench Math42.2%—
MATH Level 541.6%—

Knowledge Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 30.6 (#226), Phi 3 Mini 4k Instruct June 2024: 28.6 (#244)

Knowledge benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Expert12251051
GPQA Diamond47.4%—
Confabulations22.8%—
Vectara Hallucination Rate4.1%—
MMLU86.3%—

Multilingual Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 39.9 (#220), Phi 3 Mini 4k Instruct June 2024: 25.9 (#282)

Multilingual benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Non-English12361013
LMArena Chinese12171033
LMArena German12511031
LMArena Japanese1150954
LMArena Korean1143880
LMArena Russian12521019
LMArena French1281—
LMArena Spanish1270—

Instruction Following Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 71.1 (#157), Phi 3 Mini 4k Instruct June 2024: 54.1 (#282)

Instruction Following benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Instruction Following12421058
LiveBench Instruction Following82.7%—

Long Context Phi 3 Mini 4k Instruct June 2024 leads

Llama-3.3-70B-Instruct: 26.4 (#295), Phi 3 Mini 4k Instruct June 2024: 31.7 (#277)

Long Context benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Longer Query12561042
Fiction.LiveBench33.3%—

Writing & Preference Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 47.6 (#207), Phi 3 Mini 4k Instruct June 2024: 29.6 (#294)

Writing & Preference benchmarks
BenchmarkLlama-3.3-70B-InstructPhi 3 Mini 4k Instruct June 2024
LMArena Text12741080
LMArena Creative Writing12501045
LMArena Multi-Turn12801049
LiveBench Language39.2%—

Frequently asked questions

Is Llama-3.3-70B-Instruct better than Phi 3 Mini 4k Instruct June 2024?

Llama-3.3-70B-Instruct and Phi 3 Mini 4k Instruct June 2024 score almost the same on the Noometry Index (30.6 vs 31.3), so choose on price, context window or the category you care about most.

Is Llama-3.3-70B-Instruct or Phi 3 Mini 4k Instruct June 2024 better for coding?

They score almost the same on coding (31.0 vs 31.8); test both on your own repository before choosing.

How many benchmarks do Llama-3.3-70B-Instruct and Phi 3 Mini 4k Instruct June 2024 share?

15 benchmarks have published results for both models. Llama-3.3-70B-Instruct has 43 scored results on Noometry and Phi 3 Mini 4k Instruct June 2024 has 15.

Related comparisons

Go deeper