Model comparison
phi-3-medium 14B vs Qwen2.5 7B Instruct
phi-3-medium 14B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.7 vs 29.0), so choose on price, context window or the category you care about most.
Last verified . 5 shared benchmarks.
Summary
- They share 5 benchmarks with published results for both. phi-3-medium 14B scores higher in 2 categories and Qwen2.5 7B Instruct in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in math, where phi-3-medium 14B leads 27.3 to 12.6.
- The biggest single-benchmark swing is GPQA Diamond: 27.6% for phi-3-medium 14B and 35.5% for Qwen2.5 7B Instruct.
Side by side
| phi-3-medium 14B | Qwen2.5 7B Instruct | |
|---|---|---|
| Provider | Microsoft | Alibaba (Qwen) |
| Noometry Index | 29.7 | 29.0 |
| Released | 2024-04-23 | 2024-09 |
| Weights | Open | Open |
| Context window | — | 131K |
| Max output | — | 8K |
| Input $ / M tokens | — | $0.17 |
| Output $ / M tokens | — | $0.70 |
| Results tracked | 13 | 15 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Too close to call
phi-3-medium 14B: 36.8 (#201), Qwen2.5 7B Instruct: 36.5 (#208)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| BigCodeBench Instruct | 37.6% | 37.6% |
| BigCodeBench Complete | 48.7% | 46.1% |
Agentic & Tool Use Not comparable
phi-3-medium 14B: —, Qwen2.5 7B Instruct: 23.8 (#124)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| BALROG | — | 7.8% |
Reasoning Not comparable
phi-3-medium 14B: —, Qwen2.5 7B Instruct: 14.8 (#322)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| Epoch Capabilities Index | 121.23 | 118.51 |
| Chess Puzzles | — | 0% |
| DTBench | — | 47.7% |
| LMCA | — | 6.4% |
| Adversarial NLI | 55.8% | — |
| BIG-Bench Hard | 81.4% | — |
| HellaSwag | 82.4% | — |
| WinoGrande | 81.5% | — |
Math phi-3-medium 14B leads
phi-3-medium 14B: 27.3 (#250), Qwen2.5 7B Instruct: 12.6 (#306)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| OTIS Mock AIME 2024-2025 | — | 2.5% |
| Omni-MATH | — | 29.4% |
| MATH Level 5 | 17.6% | — |
Knowledge Qwen2.5 7B Instruct leads
phi-3-medium 14B: 9.1 (#306), Qwen2.5 7B Instruct: 17.0 (#286)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| GPQA Diamond | 27.6% | 35.5% |
| MMLU | 78% | 72.9% |
| MMLU-Pro | — | 53.9% |
| GPQA (HELM) | — | 34.1% |
| ARC (AI2) Challenge | 91.6% | — |
| OpenBookQA | 87.4% | — |
| TriviaQA | 73.9% | — |
Instruction Following Not comparable
phi-3-medium 14B: —, Qwen2.5 7B Instruct: 63.2 (#231)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| IFEval | — | 74.1% |
Writing & Preference Not comparable
phi-3-medium 14B: —, Qwen2.5 7B Instruct: 48.8 (#195)
| Benchmark | phi-3-medium 14B | Qwen2.5 7B Instruct |
|---|---|---|
| WildBench | — | 73.1% |
Frequently asked questions
Is phi-3-medium 14B better than Qwen2.5 7B Instruct?
phi-3-medium 14B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.7 vs 29.0), so choose on price, context window or the category you care about most.
Is phi-3-medium 14B or Qwen2.5 7B Instruct better for coding?
They score almost the same on coding (36.8 vs 36.5); test both on your own repository before choosing.
How many benchmarks do phi-3-medium 14B and Qwen2.5 7B Instruct share?
5 benchmarks have published results for both models. phi-3-medium 14B has 13 scored results on Noometry and Qwen2.5 7B Instruct has 15.