Alibaba (Qwen), open weights

# Qwen2.5 7B Instruct

> Qwen2.5 7B Instruct by Alibaba (Qwen), released September 2024. Ranked #320 of 354 with a Noometry Index of 29.0. API: $0.17 in / $0.70 out per M tokens. 131K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/qwen2-5-7b-instruct
- Last updated: 2026-10-10
- Title: Qwen2.5 7B Instruct Benchmarks, Price & Rank (October 2026)

Qwen2.5 7B Instruct by Alibaba (Qwen) ranks 320th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is agentic & tool use, where it ranks 124th. API pricing starts at $0.17 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #320 of 354
- **Index score:** 29.0
- **Evidence:** Confirmed 15 results
- **Provider:** [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba)
- **Released:** September 1, 2024
- **Weights:** Open weights
- **Reasoning:** No
- **Context window:** 131K
- **Max output:** 8K
- **Input price:** $0.17 / M
- **Output price:** $0.70 / M
- **Blended price:** $0.31 / M
- **Output speed:** Not measured
- **Value:** #65 of 219
- **Knowledge cutoff:** April 2024
- **Input:** text
- **Hugging Face:** [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct)

## Category scores

Each category score combines every public result we have in that category.

Qwen2.5 7B Instruct category scores

1.  Coding 36.5
2.  Agentic & Tool Use 23.8
3.  Reasoning 14.8
4.  Math 12.6
5.  Knowledge 17.0
6.  Instruction Following 63.2
7.  Writing & Preference 48.8
8.  020406080

Qwen2.5 7B Instruct category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 36.5 | #208 | 2 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 23.8 | #124 | 1 |
| [Reasoning](https://noometry.com/best/reasoning) | 14.8 | #322 | 3 |
| [Math](https://noometry.com/best/math) | 12.6 | #306 | 2 |
| [Knowledge](https://noometry.com/best/knowledge) | 17.0 | #286 | 3 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 63.2 | #231 | 1 |
| [Writing & Preference](https://noometry.com/best/writing) | 48.8 | #195 | 1 |

## Strengths and weaknesses

Categories where Qwen2.5 7B Instruct places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Qwen2.5 7B Instruct: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 36.5 | −2.2 | #208 of 340, top 62% |
| [Writing & Preference](https://noometry.com/best/writing) | 48.8 | −5.0 | #195 of 312, top 63% |
| [Instruction Following](https://noometry.com/best/instruction-following) | 63.2 | −8.0 | #231 of 305, top 76% |

### Weakest categories

Qwen2.5 7B Instruct: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Math](https://noometry.com/best/math) | 12.6 | −24.0 | #306 of 327, top 94% |
| [Reasoning](https://noometry.com/best/reasoning) | 14.8 | −8.8 | #322 of 350, top 92% |
| [Knowledge](https://noometry.com/best/knowledge) | 17.0 | −20.3 | #286 of 314, top 92% |

## Closest competitors

The models ranked just above and below Qwen2.5 7B Instruct. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen2.5 7B Instruct
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [GPT-4](https://noometry.com/models/gpt-4) | #316 | 29.1 | $37.50 | — | [Compare](https://noometry.com/compare/gpt-4-vs-qwen2-5-7b-instruct) |
| [Llama 2-7B](https://noometry.com/models/llama-2-7b) | #317 | 29.1 | — | — | [Compare](https://noometry.com/compare/llama-2-7b-vs-qwen2-5-7b-instruct) |
| [Granite 4.0 Micro](https://noometry.com/models/granite-4-0-micro) | #318 | 29.0 | $0.0408 | — | [Compare](https://noometry.com/compare/granite-4-0-micro-vs-qwen2-5-7b-instruct) |
| [Claude 3 Sonnet](https://noometry.com/models/claude-3-sonnet) | #319 | 29.0 | — | — | [Compare](https://noometry.com/compare/claude-3-sonnet-vs-qwen2-5-7b-instruct) |
| [Llama 3.2 3B](https://noometry.com/models/llama-3-2-3b) | #321 | 28.9 | $0.12 | — | [Compare](https://noometry.com/compare/llama-3-2-3b-vs-qwen2-5-7b-instruct) |
| [Qwen1.5 4b Chat](https://noometry.com/models/qwen1-5-4b-chat) | #322 | 28.8 | — | — | [Compare](https://noometry.com/compare/qwen1-5-4b-chat-vs-qwen2-5-7b-instruct) |
| [Llama 3-70B](https://noometry.com/models/llama-3-70b) | #323 | 28.8 | — | 104 | [Compare](https://noometry.com/compare/llama-3-70b-vs-qwen2-5-7b-instruct) |
| [GPT-4o](https://noometry.com/models/gpt-4o) | #324 | 28.6 | $4.38 | — | [Compare](https://noometry.com/compare/gpt-4o-vs-qwen2-5-7b-instruct) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Qwen2.5 7B Instruct Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [BigCodeBench Instruct](https://noometry.com/benchmarks/bigcodebench-instruct) | 37.6% | #41 of 64, top 65% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| [BigCodeBench Complete](https://noometry.com/benchmarks/bigcodebench-complete) | 46.1% | #42 of 66, top 64% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |

### Agentic & Tool Use

Qwen2.5 7B Instruct Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [BALROG](https://noometry.com/benchmarks/balrog) | 7.8% | #34 of 35, top 98% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

Qwen2.5 7B Instruct Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 0% | #127 of 129, top 99% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-30 |
| [DTBench](https://noometry.com/benchmarks/dtbench) | 47.7% | #143 of 151, top 95% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMCA](https://noometry.com/benchmarks/lmca) | 6.4% | #119 of 125, top 96% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 118.51 | #175 of 213, top 83% |  | [Epoch AI](https://epoch.ai/eci) | 2024-09-19 |

### Math

Qwen2.5 7B Instruct Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 2.5% | #159 of 173, top 92% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-30 |
| [Omni-MATH](https://noometry.com/benchmarks/omni-math) | 29.4% | #39 of 57, top 69% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |

### Knowledge

Qwen2.5 7B Instruct Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 35.5% | #158 of 186, top 85% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-30 |
| [MMLU-Pro](https://noometry.com/benchmarks/mmlu-pro) | 53.9% | #49 of 58, top 85% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [GPQA (HELM)](https://noometry.com/benchmarks/helm-gpqa) | 34.1% | #49 of 57, top 86% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 72.9% | #39 of 81, top 49% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Instruction Following

Qwen2.5 7B Instruct Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [IFEval](https://noometry.com/benchmarks/ifeval) | 74.1% | #52 of 57, top 92% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |

### Writing & Preference

Qwen2.5 7B Instruct Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [WildBench](https://noometry.com/benchmarks/wildbench) | 73.1% | #50 of 57, top 88% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |

## API pricing by provider

Qwen2.5 7B Instruct API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [alibaba](https://www.alibabacloud.com/help/en/model-studio/models) | $0.17 | $0.70 | — | 2026-10-10 |
| [openrouter](https://openrouter.ai/qwen/qwen-2.5-7b-instruct) | $0.10 | $0.20 | — | 2026-10-10 |
| [together](https://docs.together.ai/docs/serverless-models) | $0.30 | $0.30 | — | 2026-10-10 |

[All Alibaba (Qwen) API prices →](https://noometry.com/llm-pricing/alibaba) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare Qwen2.5 7B Instruct

-   [Qwen2.5 7B Instruct vs Claude 3 Sonnet](https://noometry.com/compare/claude-3-sonnet-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Llama 3.2 3B](https://noometry.com/compare/llama-3-2-3b-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Granite 4.0 Micro](https://noometry.com/compare/granite-4-0-micro-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Qwen1.5 4b Chat](https://noometry.com/compare/qwen1-5-4b-chat-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Llama 2-7B](https://noometry.com/compare/llama-2-7b-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Llama 3-70B](https://noometry.com/compare/llama-3-70b-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Kimi K3](https://noometry.com/compare/kimi-k3-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Grok 4.6](https://noometry.com/compare/grok-4-6-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs Muse Spark 1.3](https://noometry.com/compare/muse-spark-1-3-vs-qwen2-5-7b-instruct)
-   [Qwen2.5 7B Instruct vs DeepSeek V4 Pro](https://noometry.com/compare/deepseek-v4-pro-vs-qwen2-5-7b-instruct)

## Other Alibaba (Qwen) models

-   [Qwen3.8 Max](https://noometry.com/models/qwen3-8-max)56.8
-   [Qwen3.7 Max](https://noometry.com/models/qwen3-7-max)51.5
-   [Qwen3.6 Max Preview](https://noometry.com/models/qwen3-6-max-preview)51.5
-   [Qwen3.6 Plus](https://noometry.com/models/qwen3-6-plus)47.5
-   [Qwen3.5 397B-A17B](https://noometry.com/models/qwen3-5-397b-a17b)46.0
-   [Qwen3.8 27B](https://noometry.com/models/qwen3-8-27b)46.0
-   [Qwen3.5 Max Preview](https://noometry.com/models/qwen3-5-max-preview)45.3
-   [Qwen3.7 Plus](https://noometry.com/models/qwen3-7-plus)45.3

## Frequently asked questions

### How good is Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct by Alibaba (Qwen) ranks 320th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is agentic & tool use, where it ranks 124th. API pricing starts at $0.17 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.

### How much does Qwen2.5 7B Instruct cost?

Qwen2.5 7B Instruct costs $0.17 per million input tokens and $0.70 per million output tokens on Alibaba (Qwen)'s own API.

### What is Qwen2.5 7B Instruct's context window?

Qwen2.5 7B Instruct accepts up to 131K tokens of input and can write up to 8K tokens in one response.

### Is Qwen2.5 7B Instruct open source?

Yes. Qwen2.5 7B Instruct's weights are downloadable from Hugging Face (Qwen/Qwen2.5-7B-Instruct); check the license for commercial terms.

### What are Qwen2.5 7B Instruct's strengths and weaknesses?

Relative to other ranked models, Qwen2.5 7B Instruct places best in coding, writing & preference, instruction following and lowest in math, reasoning, knowledge.

### What is Qwen2.5 7B Instruct best at?

Its best category is agentic & tool use, where it ranks 124th on Noometry.

### Cite this page

Noometry. (2026). Qwen2.5 7B Instruct benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/qwen2-5-7b-instruct

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/qwen2-5-7b-instruct.md).
