Microsoft, open weights

# Phi-4

> Phi-4 by Microsoft, released December 2024. Ranked #279 of 354 with a Noometry Index of 31.2. API: $0.07 in / $0.14 out per M tokens. 128K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/phi-4
- Last updated: 2026-10-10
- Title: Phi-4 Benchmarks, Price & Rank (October 2026) | Noometry

Phi-4 by Microsoft ranks 279th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.2. Its strongest category is agentic & tool use, where it ranks 128th. API pricing starts at $0.07 per million input tokens and $0.14 per million output tokens, with a 128K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #279 of 354
- **Index score:** 31.2
- **Evidence:** Confirmed 37 results
- **Provider:** [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft)
- **Released:** December 11, 2024
- **Weights:** Open weights
- **Reasoning:** No
- **Context window:** 128K
- **Max output:** 4K
- **Input price:** $0.07 / M
- **Output price:** $0.14 / M
- **Blended price:** $0.0875 / M
- **Output speed:** Not measured
- **Value:** #15 of 219
- **Knowledge cutoff:** October 2023
- **Input:** text
- **Hugging Face:** [microsoft/phi-4](https://huggingface.co/microsoft/phi-4)

## Category scores

Each category score combines every public result we have in that category.

Phi-4 category scores

1.  Coding 34.4
2.  Agentic & Tool Use 22.8
3.  Reasoning 17.7
4.  Math 20.8
5.  Knowledge 32.6
6.  Multilingual 37.2
7.  Instruction Following 60.4
8.  Long Context 36.9
9.  Writing & Preference 40.5
10.  020406080

Phi-4 category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 34.4 | #239 | 4 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 22.8 | #128 | 2 |
| [Reasoning](https://noometry.com/best/reasoning) | 17.7 | #291 | 4 |
| [Math](https://noometry.com/best/math) | 20.8 | #285 | 4 |
| [Knowledge](https://noometry.com/best/knowledge) | 32.6 | #209 | 4 |
| [Multilingual](https://noometry.com/best/multilingual) | 37.2 | #237 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 60.4 | #251 | 2 |
| [Long Context](https://noometry.com/best/long-context) | 36.9 | #226 | 1 |
| [Writing & Preference](https://noometry.com/best/writing) | 40.5 | #244 | 5 |

## Strengths and weaknesses

Categories where Phi-4 places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Phi-4: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Knowledge](https://noometry.com/best/knowledge) | 32.6 | −4.7 | #209 of 314, top 67% |
| [Coding](https://noometry.com/best/coding) | 34.4 | −4.4 | #239 of 340, top 71% |
| [Long Context](https://noometry.com/best/long-context) | 36.9 | −4.0 | #226 of 296, top 77% |

### Weakest categories

Phi-4: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Math](https://noometry.com/best/math) | 20.8 | −15.8 | #285 of 327, top 88% |
| [Reasoning](https://noometry.com/best/reasoning) | 17.7 | −5.9 | #291 of 350, top 84% |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 22.8 | −7.6 | #128 of 154, top 84% |

## Closest competitors

The models ranked just above and below Phi-4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Phi-4
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Qwen-14B](https://noometry.com/models/qwen-14b) | #275 | 31.4 | — | — | [Compare](https://noometry.com/compare/phi-4-vs-qwen-14b) |
| [Phi 3 Mini 4k Instruct June 2024](https://noometry.com/models/phi-3-mini-4k-instruct-june) | #276 | 31.3 | — | — | [Compare](https://noometry.com/compare/phi-3-mini-4k-instruct-june-vs-phi-4) |
| [Gemma 1.1 7b IT](https://noometry.com/models/gemma-1-1-7b-it) | #277 | 31.3 | — | — | [Compare](https://noometry.com/compare/gemma-1-1-7b-it-vs-phi-4) |
| [Mistral Small 3](https://noometry.com/models/mistral-small-3) | #278 | 31.2 | $0.0575 | — | [Compare](https://noometry.com/compare/mistral-small-3-vs-phi-4) |
| [Mistral Small 3.2](https://noometry.com/models/mistral-small-3-2) | #280 | 31.2 | $0.13 | 68 | [Compare](https://noometry.com/compare/mistral-small-3-2-vs-phi-4) |
| [Amazon Nova Pro](https://noometry.com/models/amazon-nova-pro) | #281 | 31.0 | $1.40 | — | [Compare](https://noometry.com/compare/amazon-nova-pro-vs-phi-4) |
| [Llama 4 Maverick](https://noometry.com/models/llama-4-maverick) | #282 | 30.9 | $0.30 | 456 | [Compare](https://noometry.com/compare/llama-4-maverick-vs-phi-4) |
| [Phi-4 Mini](https://noometry.com/models/phi-4-mini) | #283 | 30.9 | $0.13 | — | [Compare](https://noometry.com/compare/phi-4-vs-phi-4-mini) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Phi-4 Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [BigCodeBench Instruct](https://noometry.com/benchmarks/bigcodebench-instruct) | 45.5% | #19 of 64, top 30% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-12-13 |
| [LiveBench Coding](https://noometry.com/benchmarks/livebench-coding) | 30.7% | #31 of 39, top 80% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1231 | #230 of 294, top 79% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [BigCodeBench Complete](https://noometry.com/benchmarks/bigcodebench-complete) | 55.4% | #17 of 66, top 26% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-12-13 |

### Agentic & Tool Use

Phi-4 Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Berkeley Function Calling Leaderboard](https://noometry.com/benchmarks/bfcl) | 28.8% | #35 of 49, top 72% | prompt | [Berkeley Function Calling Leaderboard](https://gorilla.cs.berkeley.edu/leaderboard.html) |  |
| [BALROG](https://noometry.com/benchmarks/balrog) | 11.6% | #32 of 35, top 92% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

Phi-4 Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 1% | #107 of 129, top 83% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |
| [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning) | 47.8% | #21 of 39, top 54% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1220 | #231 of 297, top 78% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LiveBench Data Analysis](https://noometry.com/benchmarks/livebench-data-analysis) | 45.2% | #30 of 39, top 77% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 130.42 | #137 of 213, top 65% |  | [Epoch AI](https://epoch.ai/eci) | 2024-12-12 |
| [LiveBench](https://noometry.com/benchmarks/livebench) | 41.6% | #29 of 39, top 75% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Math

Phi-4 Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 13.8% | #131 of 173, top 76% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-02-25 |
| [LiveBench Math](https://noometry.com/benchmarks/livebench-math) | 42% | #25 of 39, top 65% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1246 | #216 of 285, top 76% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 64.9% | #35 of 79, top 45% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-31 |

### Knowledge

Phi-4 Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 56.1% | #120 of 186, top 65% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-31 |
| [Confabulations](https://noometry.com/benchmarks/confabulations) (lower is better) | 29.4% | #46 of 51, top 91% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/confabulations) |  |
| [Vectara Hallucination Rate](https://noometry.com/benchmarks/vectara-hallucination) (lower is better) | 3.7% | #3 of 96, top 4% |  | [Vectara Hallucination Leaderboard](https://github.com/vectara/hallucination-leaderboard) |  |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1203 | #220 of 273, top 81% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 84.8% | #8 of 81, top 10% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Multilingual

Phi-4 Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1197 | #237 of 297, top 80% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1212 | #228 of 285, top 80% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1224 | #190 of 223, top 86% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1222 | #185 of 231, top 81% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1158 | #172 of 211, top 82% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1151 | #177 of 213, top 84% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1209 | #231 of 283, top 82% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1234 | #188 of 226, top 84% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

Phi-4 Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LiveBench Instruction Following](https://noometry.com/benchmarks/livebench-if) | 58.4% | #29 of 39, top 75% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1201 | #234 of 298, top 79% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

Phi-4 Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1217 | #236 of 291, top 82% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

Phi-4 Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1217 | #240 of 297, top 81% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1182 | #239 of 295, top 82% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Short-Story Creative Writing](https://noometry.com/benchmarks/lech-mazur-writing) | 62.6% | #36 of 39, top 93% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1206 | #236 of 295, top 80% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LiveBench Language](https://noometry.com/benchmarks/livebench-language) | 25.6% | #31 of 39, top 80% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## API pricing by provider

Phi-4 API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [azure](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) | $0.13 | $0.50 | — | 2026-10-10 |
| [openrouter](https://openrouter.ai/microsoft/phi-4) | $0.07 | $0.14 | — | 2026-10-10 |

[All Microsoft API prices →](https://noometry.com/llm-pricing/microsoft) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare Phi-4

-   [Phi-4 vs Mistral Small 3](https://noometry.com/compare/mistral-small-3-vs-phi-4)
-   [Phi-4 vs Mistral Small 3.2](https://noometry.com/compare/mistral-small-3-2-vs-phi-4)
-   [Phi-4 vs Gemma 1.1 7b IT](https://noometry.com/compare/gemma-1-1-7b-it-vs-phi-4)
-   [Phi-4 vs Amazon Nova Pro](https://noometry.com/compare/amazon-nova-pro-vs-phi-4)
-   [Phi-4 vs Phi 3 Mini 4k Instruct June 2024](https://noometry.com/compare/phi-3-mini-4k-instruct-june-vs-phi-4)
-   [Phi-4 vs Llama 4 Maverick](https://noometry.com/compare/llama-4-maverick-vs-phi-4)
-   [Phi-4 vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-phi-4)
-   [Phi-4 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-phi-4)
-   [Phi-4 vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-phi-4)
-   [Phi-4 vs Kimi K3](https://noometry.com/compare/kimi-k3-vs-phi-4)
-   [Phi-4 vs Grok 4.6](https://noometry.com/compare/grok-4-6-vs-phi-4)
-   [Phi-4 vs Qwen3.8 Max](https://noometry.com/compare/phi-4-vs-qwen3-8-max)
-   [Phi-4 vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-phi-4)
-   [Phi-4 vs Muse Spark 1.3](https://noometry.com/compare/muse-spark-1-3-vs-phi-4)

## Other Microsoft models

-   [Wizardlm 70b](https://noometry.com/models/wizardlm-70b)33.0
-   [Phi 3 Medium 4k Instruct](https://noometry.com/models/phi-3-medium-4k-instruct)33.0
-   [Wizardlm 13b](https://noometry.com/models/wizardlm-13b)31.4
-   [Phi 3 Mini 4k Instruct June 2024](https://noometry.com/models/phi-3-mini-4k-instruct-june)31.3
-   [Phi-4 Mini](https://noometry.com/models/phi-4-mini)30.9
-   [Phi 3 Mini 128k Instruct](https://noometry.com/models/phi-3-mini-128k-instruct)29.7
-   [phi-3-medium 14B](https://noometry.com/models/phi-3-medium-14b)29.7
-   [Phi 3 Small 8k Instruct](https://noometry.com/models/phi-3-small-8k-instruct)29.3

## Frequently asked questions

### How good is Phi-4?

Phi-4 by Microsoft ranks 279th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.2. Its strongest category is agentic & tool use, where it ranks 128th. API pricing starts at $0.07 per million input tokens and $0.14 per million output tokens, with a 128K-token context window.

### How much does Phi-4 cost?

Phi-4 costs $0.07 per million input tokens and $0.14 per million output tokens on openrouter.

### What is Phi-4's context window?

Phi-4 accepts up to 128K tokens of input and can write up to 4K tokens in one response.

### Is Phi-4 open source?

Yes. Phi-4's weights are downloadable from Hugging Face (microsoft/phi-4); check the license for commercial terms.

### What are Phi-4's strengths and weaknesses?

Relative to other ranked models, Phi-4 places best in knowledge, coding, long context and lowest in math, reasoning, agentic & tool use.

### What is Phi-4 best at?

Its best category is agentic & tool use, where it ranks 128th on Noometry.

### Cite this page

Noometry. (2026). Phi-4 benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/phi-4

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/phi-4.md).
