Model comparison

Llama 3.1-405B vs Phi 3 Mini 4k Instruct June 2024

Llama 3.1-405B and Phi 3 Mini 4k Instruct June 2024 score almost the same on the Noometry Index (30.7 vs 31.3), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Llama 3.1-405B scores higher in 6 categories and Phi 3 Mini 4k Instruct June 2024 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Llama 3.1-405B leads 40.7 to 25.9.

Side by side

Llama 3.1-405B and Phi 3 Mini 4k Instruct June 2024 specifications
Llama 3.1-405BPhi 3 Mini 4k Instruct June 2024
ProviderMetaMicrosoft
Noometry Index30.731.3
Released2024-07-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Llama 3.1-405B: 33.1 (#262), Phi 3 Mini 4k Instruct June 2024: 31.8 (#279)

Coding benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Coding12911093
WeirdML21.4%—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), Phi 3 Mini 4k Instruct June 2024: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Phi 3 Mini 4k Instruct June 2024 leads

Llama 3.1-405B: 16.8 (#300), Phi 3 Mini 4k Instruct June 2024: 20.8 (#231)

Reasoning benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Hard Prompts12691087
SimpleBench23%—
Kagi LLM Benchmark45%—
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Phi 3 Mini 4k Instruct June 2024 leads

Llama 3.1-405B: 18.4 (#290), Phi 3 Mini 4k Instruct June 2024: 33.0 (#208)

Math benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Math12811152
OTIS Mock AIME 2024-20259.7%—
Omni-MATH24.9%—
MATH Level 549.8%—

Knowledge Llama 3.1-405B leads

Llama 3.1-405B: 30.4 (#227), Phi 3 Mini 4k Instruct June 2024: 28.6 (#244)

Knowledge benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Expert12431051
GPQA Diamond50.9%—
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Llama 3.1-405B leads

Llama 3.1-405B: 40.7 (#214), Phi 3 Mini 4k Instruct June 2024: 25.9 (#282)

Multilingual benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Non-English12481013
LMArena Chinese12421033
LMArena German12521031
LMArena Japanese1208954
LMArena Korean1184880
LMArena Russian12651019
LMArena French1279—
LMArena Spanish1260—

Instruction Following Llama 3.1-405B leads

Llama 3.1-405B: 65.9 (#214), Phi 3 Mini 4k Instruct June 2024: 54.1 (#282)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Instruction Following12591058
IFEval81.1%—

Long Context Llama 3.1-405B leads

Llama 3.1-405B: 38.4 (#197), Phi 3 Mini 4k Instruct June 2024: 31.7 (#277)

Long Context benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Longer Query12661042

Writing & Preference Llama 3.1-405B leads

Llama 3.1-405B: 38.9 (#251), Phi 3 Mini 4k Instruct June 2024: 29.6 (#294)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BPhi 3 Mini 4k Instruct June 2024
LMArena Text12841080
LMArena Creative Writing12621045
LMArena Multi-Turn12971049
EQ-Bench Creative Writing870—
WildBench78.3%—

Frequently asked questions

Is Llama 3.1-405B better than Phi 3 Mini 4k Instruct June 2024?

Llama 3.1-405B and Phi 3 Mini 4k Instruct June 2024 score almost the same on the Noometry Index (30.7 vs 31.3), so choose on price, context window or the category you care about most.

Is Llama 3.1-405B or Phi 3 Mini 4k Instruct June 2024 better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 31.8 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Phi 3 Mini 4k Instruct June 2024 share?

15 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Phi 3 Mini 4k Instruct June 2024 has 15.

Related comparisons

Go deeper