Model comparison

phi-3-medium 14B vs Qwen2.5 7B Instruct

phi-3-medium 14B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.7 vs 29.0), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 5 benchmarks with published results for both. phi-3-medium 14B scores higher in 2 categories and Qwen2.5 7B Instruct in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where phi-3-medium 14B leads 27.3 to 12.6.
  • The biggest single-benchmark swing is GPQA Diamond: 27.6% for phi-3-medium 14B and 35.5% for Qwen2.5 7B Instruct.

Side by side

phi-3-medium 14B and Qwen2.5 7B Instruct specifications
phi-3-medium 14BQwen2.5 7B Instruct
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.729.0
Released2024-04-232024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.17
Output $ / M tokens—$0.70
Results tracked1315

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

phi-3-medium 14B: 36.8 (#201), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
BigCodeBench Instruct37.6%37.6%
BigCodeBench Complete48.7%46.1%

Agentic & Tool Use Not comparable

phi-3-medium 14B: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
BALROG—7.8%

Reasoning Not comparable

phi-3-medium 14B: —, Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
Epoch Capabilities Index121.23118.51
Chess Puzzles—0%
DTBench—47.7%
LMCA—6.4%
Adversarial NLI55.8%—
BIG-Bench Hard81.4%—
HellaSwag82.4%—
WinoGrande81.5%—

Math phi-3-medium 14B leads

phi-3-medium 14B: 27.3 (#250), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%
MATH Level 517.6%—

Knowledge Qwen2.5 7B Instruct leads

phi-3-medium 14B: 9.1 (#306), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
GPQA Diamond27.6%35.5%
MMLU78%72.9%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
ARC (AI2) Challenge91.6%—
OpenBookQA87.4%—
TriviaQA73.9%—

Instruction Following Not comparable

phi-3-medium 14B: —, Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
IFEval—74.1%

Writing & Preference Not comparable

phi-3-medium 14B: —, Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
Benchmarkphi-3-medium 14BQwen2.5 7B Instruct
WildBench—73.1%

Frequently asked questions

Is phi-3-medium 14B better than Qwen2.5 7B Instruct?

phi-3-medium 14B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.7 vs 29.0), so choose on price, context window or the category you care about most.

Is phi-3-medium 14B or Qwen2.5 7B Instruct better for coding?

They score almost the same on coding (36.8 vs 36.5); test both on your own repository before choosing.

How many benchmarks do phi-3-medium 14B and Qwen2.5 7B Instruct share?

5 benchmarks have published results for both models. phi-3-medium 14B has 13 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper