Model comparison

Llama 2-34B vs Phi-3.5-mini

Neither Llama 2-34B nor Phi-3.5-mini has enough public benchmark results to be ranked yet; the rows below show what has been published.

Last verified . 3 shared benchmarks.

Llama 2-34B Meta

—

Unranked

Phi-3.5-mini Microsoft

36.9

Unranked Sparse

Summary

  • They share 3 benchmarks with published results for both.
  • Phi-3.5-mini has downloadable open weights; the other is API-only.

Side by side

Llama 2-34B and Phi-3.5-mini specifications
Llama 2-34BPhi-3.5-mini
ProviderMetaMicrosoft
Noometry Index—36.9
Released2023-07-182024-08-16
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked105

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 2-34B: —, Phi-3.5-mini: 34.6 (#232)

Coding benchmarks
BenchmarkLlama 2-34BPhi-3.5-mini
BigCodeBench Instruct—32.8%
BigCodeBench Complete—38.5%

Reasoning Not comparable

Llama 2-34B: —, Phi-3.5-mini: —

Reasoning benchmarks
BenchmarkLlama 2-34BPhi-3.5-mini
PIQA81.9%81%
BIG-Bench Hard44.1%—
Epoch Capabilities Index105.2—
WinoGrande76.7%—

Math Not comparable

Llama 2-34B: —, Phi-3.5-mini: —

Math benchmarks
BenchmarkLlama 2-34BPhi-3.5-mini
GSM8K42.2%86.2%

Knowledge Not comparable

Llama 2-34B: —, Phi-3.5-mini: —

Knowledge benchmarks
BenchmarkLlama 2-34BPhi-3.5-mini
BoolQ83.7%78%
ARC (AI2) Challenge54.5%—
MMLU62.6%—
OpenBookQA58.2%—
TriviaQA84.6%—

Frequently asked questions

Is Llama 2-34B better than Phi-3.5-mini?

Neither Llama 2-34B nor Phi-3.5-mini has enough public benchmark results to be ranked yet; the rows below show what has been published.

How many benchmarks do Llama 2-34B and Phi-3.5-mini share?

3 benchmarks have published results for both models. Llama 2-34B has 10 scored results on Noometry and Phi-3.5-mini has 5.

Related comparisons

Go deeper