Alibaba (Qwen), open weights

Qwen2.5 7B Instruct

Qwen2.5 7B Instruct by Alibaba (Qwen) ranks 320th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is agentic & tool use, where it ranks 124th. API pricing starts at $0.17 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#320 of 354
Index score
29.0
Evidence
Confirmed 15 results
Released
September 1, 2024
Weights
Open weights
Reasoning
No
Context window
131K
Max output
8K
Input price
$0.17 / M
Output price
$0.70 / M
Blended price
$0.31 / M
Output speed
Not measured
Value
#65 of 219
Knowledge cutoff
April 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

Qwen2.5 7B Instruct category scores
  1. Coding 36.5
  2. Agentic & Tool Use 23.8
  3. Reasoning 14.8
  4. Math 12.6
  5. Knowledge 17.0
  6. Instruction Following 63.2
  7. Writing & Preference 48.8
Qwen2.5 7B Instruct category ranks
CategoryScoreRankResults
Coding36.5#2082
Agentic & Tool Use23.8#1241
Reasoning14.8#3223
Math12.6#3062
Knowledge17.0#2863
Instruction Following63.2#2311
Writing & Preference48.8#1951

Strengths and weaknesses

Categories where Qwen2.5 7B Instruct places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Qwen2.5 7B Instruct: strongest categories
CategoryScorevs medianRank
Coding36.5−2.2#208 of 340, top 62%
Writing & Preference48.8−5.0#195 of 312, top 63%
Instruction Following63.2−8.0#231 of 305, top 76%

Weakest categories

Qwen2.5 7B Instruct: weakest categories
CategoryScorevs medianRank
Math12.6−24.0#306 of 327, top 94%
Reasoning14.8−8.8#322 of 350, top 92%
Knowledge17.0−20.3#286 of 314, top 92%

Closest competitors

The models ranked just above and below Qwen2.5 7B Instruct. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen2.5 7B Instruct
ModelRankScoreBlended $/MSpeed
GPT-4#31629.1$37.50—Compare
Llama 2-7B#31729.1——Compare
Granite 4.0 Micro#31829.0$0.0408—Compare
Claude 3 Sonnet#31929.0——Compare
Llama 3.2 3B#32128.9$0.12—Compare
Qwen1.5 4b Chat#32228.8——Compare
Llama 3-70B#32328.8—104Compare
GPT-4o#32428.6$4.38—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Qwen2.5 7B Instruct Coding benchmark results
BenchmarkScorePositionSettingSourceDate
BigCodeBench Instruct37.6%#41 of 64, top 65%BigCodeBench2024-09-19
BigCodeBench Complete46.1%#42 of 66, top 64%BigCodeBench2024-09-19

Agentic & Tool Use

Qwen2.5 7B Instruct Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
BALROG7.8%#34 of 35, top 98%Epoch AI

Reasoning

Qwen2.5 7B Instruct Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles0%#127 of 129, top 99%Epoch AI2026-08-30
DTBench47.7%#143 of 151, top 95%Epoch AI
LMCA6.4%#119 of 125, top 96%Epoch AI
Epoch Capabilities Index118.51#175 of 213, top 83%Epoch AI2024-09-19

Math

Qwen2.5 7B Instruct Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20252.5%#159 of 173, top 92%Epoch AI2026-08-30
Omni-MATH29.4%#39 of 57, top 69%HELM Capabilities

Knowledge

Qwen2.5 7B Instruct Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond35.5%#158 of 186, top 85%Epoch AI2026-08-30
MMLU-Pro53.9%#49 of 58, top 85%HELM Capabilities
GPQA (HELM)34.1%#49 of 57, top 86%HELM Capabilities
MMLU72.9%#39 of 81, top 49%Epoch AI

Instruction Following

Qwen2.5 7B Instruct Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval74.1%#52 of 57, top 92%HELM Capabilities

Writing & Preference

Qwen2.5 7B Instruct Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
WildBench73.1%#50 of 57, top 88%HELM Capabilities

API pricing by provider

Qwen2.5 7B Instruct API prices
RouteInput $/MOutput $/MCached input $/MChecked
alibaba$0.17$0.70—2026-10-10
openrouter$0.10$0.20—2026-10-10
together$0.30$0.30—2026-10-10

Compare Qwen2.5 7B Instruct

Other Alibaba (Qwen) models

Frequently asked questions

How good is Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct by Alibaba (Qwen) ranks 320th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is agentic & tool use, where it ranks 124th. API pricing starts at $0.17 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.

How much does Qwen2.5 7B Instruct cost?

Qwen2.5 7B Instruct costs $0.17 per million input tokens and $0.70 per million output tokens on Alibaba (Qwen)'s own API.

What is Qwen2.5 7B Instruct's context window?

Qwen2.5 7B Instruct accepts up to 131K tokens of input and can write up to 8K tokens in one response.

Is Qwen2.5 7B Instruct open source?

Yes. Qwen2.5 7B Instruct's weights are downloadable from Hugging Face (Qwen/Qwen2.5-7B-Instruct); check the license for commercial terms.

What are Qwen2.5 7B Instruct's strengths and weaknesses?

Relative to other ranked models, Qwen2.5 7B Instruct places best in coding, writing & preference, instruction following and lowest in math, reasoning, knowledge.

What is Qwen2.5 7B Instruct best at?

Its best category is agentic & tool use, where it ranks 124th on Noometry.