Microsoft, open weights

# Phi-2

> Phi-2 by Microsoft, released December 2023. 11 published benchmark results. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/phi-2
- Last updated: 2026-10-10
- Title: Phi-2 Benchmarks, Price & Rank (October 2026) | Noometry

Phi-2 by Microsoft has 11 published benchmark results on Noometry, not yet enough to be ranked.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** Unranked
- **Index score:** Not ranked
- **Evidence:** 11 results
- **Provider:** [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft)
- **Released:** December 12, 2023
- **Weights:** Open weights
- **Reasoning:** Unknown
- **Context window:** —
- **Max output:** —
- **Input price:** Not listed
- **Output price:** Not listed
- **Blended price:** Not listed
- **Output speed:** Not measured
- **Value:** Not ranked
- **Knowledge cutoff:** Unknown

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Phi-2 Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [HumanEval+](https://noometry.com/benchmarks/humaneval-plus) | 45.1% | #35 of 45, top 78% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| [MBPP+](https://noometry.com/benchmarks/mbpp-plus) | 54.2% | #31 of 38, top 82% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |

### Reasoning

Phi-2 Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Adversarial NLI](https://noometry.com/benchmarks/anli) | 42.5% | #9 of 9, top 100% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BIG-Bench Hard](https://noometry.com/benchmarks/bbh) | 59.4% | #13 of 27, top 49% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 107.94 | #193 of 213, top 91% |  | [Epoch AI](https://epoch.ai/eci) | 2023-12-12 |
| [HellaSwag](https://noometry.com/benchmarks/hellaswag) | 53.6% | #28 of 29, top 97% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WinoGrande](https://noometry.com/benchmarks/winogrande) | 54.7% | #42 of 43, top 98% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Knowledge

Phi-2 Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [ARC (AI2) Challenge](https://noometry.com/benchmarks/arc-challenge) | 75.9% | #16 of 39, top 42% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 58.4% | #64 of 81, top 80% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [OpenBookQA](https://noometry.com/benchmarks/openbookqa) | 73.6% | #9 of 19, top 48% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [TriviaQA](https://noometry.com/benchmarks/triviaqa) | 45.2% | #25 of 25, top 100% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare Phi-2

-   [Phi-2 vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-phi-2)
-   [Phi-2 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-phi-2)
-   [Phi-2 vs Claude Opus 5.5](https://noometry.com/compare/claude-opus-5-5-vs-phi-2)
-   [Phi-2 vs Claude Opus 5](https://noometry.com/compare/claude-opus-5-vs-phi-2)
-   [Phi-2 vs Claude Fable 5](https://noometry.com/compare/claude-fable-5-vs-phi-2)
-   [Phi-2 vs GPT-6.1 Sol](https://noometry.com/compare/gpt-6-1-sol-vs-phi-2)
-   [Phi-2 vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-phi-2)
-   [Phi-2 vs Kimi K3](https://noometry.com/compare/kimi-k3-vs-phi-2)
-   [Phi-2 vs Grok 4.6](https://noometry.com/compare/grok-4-6-vs-phi-2)
-   [Phi-2 vs Qwen3.8 Max](https://noometry.com/compare/phi-2-vs-qwen3-8-max)
-   [Phi-2 vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-phi-2)
-   [Phi-2 vs Muse Spark 1.3](https://noometry.com/compare/muse-spark-1-3-vs-phi-2)
-   [Phi-2 vs DeepSeek V4 Pro](https://noometry.com/compare/deepseek-v4-pro-vs-phi-2)
-   [Phi-2 vs MiMo-V2.6-Pro](https://noometry.com/compare/mimo-v2-6-pro-vs-phi-2)

## Other Microsoft models

-   [Wizardlm 70b](https://noometry.com/models/wizardlm-70b)33.0
-   [Phi 3 Medium 4k Instruct](https://noometry.com/models/phi-3-medium-4k-instruct)33.0
-   [Wizardlm 13b](https://noometry.com/models/wizardlm-13b)31.4
-   [Phi 3 Mini 4k Instruct June 2024](https://noometry.com/models/phi-3-mini-4k-instruct-june)31.3
-   [Phi-4](https://noometry.com/models/phi-4)31.2
-   [Phi-4 Mini](https://noometry.com/models/phi-4-mini)30.9
-   [Phi 3 Mini 128k Instruct](https://noometry.com/models/phi-3-mini-128k-instruct)29.7
-   [phi-3-medium 14B](https://noometry.com/models/phi-3-medium-14b)29.7

## Frequently asked questions

### How good is Phi-2?

Phi-2 by Microsoft has 11 published benchmark results on Noometry, not yet enough to be ranked.

### Is Phi-2 open source?

Yes. Phi-2's weights are downloadable; check the license for commercial terms.

### Cite this page

Noometry. (2026). Phi-2 benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/phi-2

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/phi-2.md).
