Model comparison

Qwen2.5 7B Instruct vs Qwen3 8B

Qwen3 8B is the stronger model overall, scoring 33.7 to 29.0 on the Noometry Index.

Last verified . 6 shared benchmarks.

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Qwen2.5 7B Instruct scores higher in 1 category and Qwen3 8B in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3 8B leads 34.9 to 12.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.5% for Qwen2.5 7B Instruct and 56.1% for Qwen3 8B.
  • Both cost about the same: $0.17 input and $0.70 output per million tokens.

Side by side

Qwen2.5 7B Instruct and Qwen3 8B specifications
Qwen2.5 7B InstructQwen3 8B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index29.033.7
Released2024-092025-04
WeightsOpenOpen
Context window131K131K
Max output8K8K
Input $ / M tokens$0.17$0.18
Output $ / M tokens$0.70$0.70
Results tracked1511

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

Qwen2.5 7B Instruct: 36.5 (#208), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
SciCode—22.6%
BigCodeBench Instruct37.6%—
BigCodeBench Complete46.1%—

Agentic & Tool Use Qwen3 8B leads

Qwen2.5 7B Instruct: 23.8 (#124), Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
Berkeley Function Calling Leaderboard—42.6%
BALROG7.8%—

Reasoning Qwen3 8B leads

Qwen2.5 7B Instruct: 14.8 (#322), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
Chess Puzzles0%5%
DTBench47.7%59.7%
LMCA6.4%8.8%
Epoch Capabilities Index118.51136.17
CritPt—0%

Math Qwen3 8B leads

Qwen2.5 7B Instruct: 12.6 (#306), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
OTIS Mock AIME 2024-20252.5%56.1%
Omni-MATH29.4%—

Knowledge Qwen3 8B leads

Qwen2.5 7B Instruct: 17.0 (#286), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
GPQA Diamond35.5%56.8%
MMLU-Pro53.9%—
Vectara Hallucination Rate—4.8%
GPQA (HELM)34.1%—
MMLU72.9%—

Instruction Following Not comparable

Qwen2.5 7B Instruct: 63.2 (#231), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
IFEval74.1%—

Long Context Not comparable

Qwen2.5 7B Instruct: —, Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
Fiction.LiveBench—62.1%

Writing & Preference Not comparable

Qwen2.5 7B Instruct: 48.8 (#195), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkQwen2.5 7B InstructQwen3 8B
WildBench73.1%—

Frequently asked questions

Is Qwen2.5 7B Instruct better than Qwen3 8B?

Qwen3 8B is the stronger model overall, scoring 33.7 to 29.0 on the Noometry Index.

Which is cheaper, Qwen2.5 7B Instruct or Qwen3 8B?

Qwen2.5 7B Instruct is cheaper. It lists at $0.17 per million input tokens and $0.70 per million output tokens; Qwen3 8B lists at $0.18 and $0.70.

Is Qwen2.5 7B Instruct or Qwen3 8B better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Qwen2.5 7B Instruct and Qwen3 8B share?

6 benchmarks have published results for both models. Qwen2.5 7B Instruct has 15 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper