Model comparison

Llama 2-34B vs Phi-3.5-MoE

Neither Llama 2-34B nor Phi-3.5-MoE has enough public benchmark results to be ranked yet; the rows below show what has been published.

Last verified . 3 shared benchmarks.

Llama 2-34B Meta

—

Unranked

Phi-3.5-MoE Microsoft

—

Unranked

Summary

  • They share 3 benchmarks with published results for both.
  • Phi-3.5-MoE has downloadable open weights; the other is API-only.

Side by side

Llama 2-34B and Phi-3.5-MoE specifications
Llama 2-34BPhi-3.5-MoE
ProviderMetaMicrosoft
Noometry Index——
Released2023-07-182024-08-17
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked103

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Reasoning Not comparable

Llama 2-34B: —, Phi-3.5-MoE: —

Reasoning benchmarks
BenchmarkLlama 2-34BPhi-3.5-MoE
PIQA81.9%88.6%
BIG-Bench Hard44.1%—
Epoch Capabilities Index105.2—
WinoGrande76.7%—

Math Not comparable

Llama 2-34B: —, Phi-3.5-MoE: —

Math benchmarks
BenchmarkLlama 2-34BPhi-3.5-MoE
GSM8K42.2%88.7%

Knowledge Not comparable

Llama 2-34B: —, Phi-3.5-MoE: —

Knowledge benchmarks
BenchmarkLlama 2-34BPhi-3.5-MoE
BoolQ83.7%84.6%
ARC (AI2) Challenge54.5%—
MMLU62.6%—
OpenBookQA58.2%—
TriviaQA84.6%—

Frequently asked questions

Is Llama 2-34B better than Phi-3.5-MoE?

Neither Llama 2-34B nor Phi-3.5-MoE has enough public benchmark results to be ranked yet; the rows below show what has been published.

How many benchmarks do Llama 2-34B and Phi-3.5-MoE share?

3 benchmarks have published results for both models. Llama 2-34B has 10 scored results on Noometry and Phi-3.5-MoE has 3.

Related comparisons

Go deeper