Alibaba (Qwen), open weights
Qwen2.5 7B Instruct
Qwen2.5 7B Instruct by Alibaba (Qwen) ranks 320th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is agentic & tool use, where it ranks 124th. API pricing starts at $0.17 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #320 of 354
- Index score
- 29.0
- Evidence
- Confirmed 15 results
- Provider
Alibaba (Qwen)
- Released
- September 1, 2024
- Weights
- Open weights
- Reasoning
- No
- Context window
- 131K
- Max output
- 8K
- Input price
- $0.17 / M
- Output price
- $0.70 / M
- Blended price
- $0.31 / M
- Output speed
- Not measured
- Value
- #65 of 219
- Knowledge cutoff
- April 2024
- Input
- text
- Hugging Face
- Qwen/Qwen2.5-7B-Instruct
Category scores
Each category score combines every public result we have in that category.
- Coding 36.5
- Agentic & Tool Use 23.8
- Reasoning 14.8
- Math 12.6
- Knowledge 17.0
- Instruction Following 63.2
- Writing & Preference 48.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 36.5 | #208 | 2 |
| Agentic & Tool Use | 23.8 | #124 | 1 |
| Reasoning | 14.8 | #322 | 3 |
| Math | 12.6 | #306 | 2 |
| Knowledge | 17.0 | #286 | 3 |
| Instruction Following | 63.2 | #231 | 1 |
| Writing & Preference | 48.8 | #195 | 1 |
Strengths and weaknesses
Categories where Qwen2.5 7B Instruct places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 36.5 | −2.2 | #208 of 340, top 62% |
| Writing & Preference | 48.8 | −5.0 | #195 of 312, top 63% |
| Instruction Following | 63.2 | −8.0 | #231 of 305, top 76% |
Closest competitors
The models ranked just above and below Qwen2.5 7B Instruct. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-4 | #316 | 29.1 | $37.50 | — | Compare |
| Llama 2-7B | #317 | 29.1 | — | — | Compare |
| Granite 4.0 Micro | #318 | 29.0 | $0.0408 | — | Compare |
| Claude 3 Sonnet | #319 | 29.0 | — | — | Compare |
| Llama 3.2 3B | #321 | 28.9 | $0.12 | — | Compare |
| Qwen1.5 4b Chat | #322 | 28.8 | — | — | Compare |
| Llama 3-70B | #323 | 28.8 | — | 104 | Compare |
| GPT-4o | #324 | 28.6 | $4.38 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| BigCodeBench Instruct | 37.6% | #41 of 64, top 65% | BigCodeBench | 2024-09-19 | |
| BigCodeBench Complete | 46.1% | #42 of 66, top 64% | BigCodeBench | 2024-09-19 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| BALROG | 7.8% | #34 of 35, top 98% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 0% | #127 of 129, top 99% | Epoch AI | 2026-08-30 | |
| DTBench | 47.7% | #143 of 151, top 95% | Epoch AI | ||
| LMCA | 6.4% | #119 of 125, top 96% | Epoch AI | ||
| Epoch Capabilities Index | 118.51 | #175 of 213, top 83% | Epoch AI | 2024-09-19 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 2.5% | #159 of 173, top 92% | Epoch AI | 2026-08-30 | |
| Omni-MATH | 29.4% | #39 of 57, top 69% | HELM Capabilities |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 35.5% | #158 of 186, top 85% | Epoch AI | 2026-08-30 | |
| MMLU-Pro | 53.9% | #49 of 58, top 85% | HELM Capabilities | ||
| GPQA (HELM) | 34.1% | #49 of 57, top 86% | HELM Capabilities | ||
| MMLU | 72.9% | #39 of 81, top 49% | Epoch AI |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 74.1% | #52 of 57, top 92% | HELM Capabilities |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WildBench | 73.1% | #50 of 57, top 88% | HELM Capabilities |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| alibaba | $0.17 | $0.70 | — | 2026-10-10 |
| openrouter | $0.10 | $0.20 | — | 2026-10-10 |
| together | $0.30 | $0.30 | — | 2026-10-10 |
Compare Qwen2.5 7B Instruct
- Qwen2.5 7B Instruct vs Claude 3 Sonnet
- Qwen2.5 7B Instruct vs Llama 3.2 3B
- Qwen2.5 7B Instruct vs Granite 4.0 Micro
- Qwen2.5 7B Instruct vs Qwen1.5 4b Chat
- Qwen2.5 7B Instruct vs Llama 2-7B
- Qwen2.5 7B Instruct vs Llama 3-70B
- Qwen2.5 7B Instruct vs GPT-6 Astra
- Qwen2.5 7B Instruct vs Claude Fable 5.1
- Qwen2.5 7B Instruct vs Gemini 3.8 Flash
- Qwen2.5 7B Instruct vs Kimi K3
- Qwen2.5 7B Instruct vs Grok 4.6
- Qwen2.5 7B Instruct vs GLM-5.3
- Qwen2.5 7B Instruct vs Muse Spark 1.3
- Qwen2.5 7B Instruct vs DeepSeek V4 Pro
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen2.5 7B Instruct?
Qwen2.5 7B Instruct by Alibaba (Qwen) ranks 320th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is agentic & tool use, where it ranks 124th. API pricing starts at $0.17 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.
How much does Qwen2.5 7B Instruct cost?
Qwen2.5 7B Instruct costs $0.17 per million input tokens and $0.70 per million output tokens on Alibaba (Qwen)'s own API.
What is Qwen2.5 7B Instruct's context window?
Qwen2.5 7B Instruct accepts up to 131K tokens of input and can write up to 8K tokens in one response.
Is Qwen2.5 7B Instruct open source?
Yes. Qwen2.5 7B Instruct's weights are downloadable from Hugging Face (Qwen/Qwen2.5-7B-Instruct); check the license for commercial terms.
What are Qwen2.5 7B Instruct's strengths and weaknesses?
Relative to other ranked models, Qwen2.5 7B Instruct places best in coding, writing & preference, instruction following and lowest in math, reasoning, knowledge.
What is Qwen2.5 7B Instruct best at?
Its best category is agentic & tool use, where it ranks 124th on Noometry.