Model comparison

Qwen2.5-Coder (1.5B) vs Qwen3 8B

Qwen3 8B has enough public results to be ranked (#238); Qwen2.5-Coder (1.5B) does not yet, so treat this comparison as directional.

Last verified . 1 shared benchmarks.

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • They share 1 benchmark with published results for both.

Side by side

Qwen2.5-Coder (1.5B) and Qwen3 8B specifications
Qwen2.5-Coder (1.5B)Qwen3 8B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index—33.7
Released2024-09-182025-04
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked611

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Qwen2.5-Coder (1.5B): —, Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkQwen2.5-Coder (1.5B)Qwen3 8B
SciCode—22.6%

Agentic & Tool Use Not comparable

Qwen2.5-Coder (1.5B): —, Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkQwen2.5-Coder (1.5B)Qwen3 8B
Berkeley Function Calling Leaderboard—42.6%

Reasoning Not comparable

Qwen2.5-Coder (1.5B): —, Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkQwen2.5-Coder (1.5B)Qwen3 8B
Epoch Capabilities Index113.14136.17
CritPt—0%
Chess Puzzles—5%
DTBench—59.7%
LMCA—8.8%
HellaSwag76.8%—
WinoGrande72.9%—

Math Not comparable

Qwen2.5-Coder (1.5B): —, Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkQwen2.5-Coder (1.5B)Qwen3 8B
OTIS Mock AIME 2024-2025—56.1%
GSM8K86.7%—

Knowledge Not comparable

Qwen2.5-Coder (1.5B): —, Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkQwen2.5-Coder (1.5B)Qwen3 8B
GPQA Diamond—56.8%
Vectara Hallucination Rate—4.8%
ARC (AI2) Challenge60.9%—
MMLU68%—

Long Context Not comparable

Qwen2.5-Coder (1.5B): —, Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkQwen2.5-Coder (1.5B)Qwen3 8B
Fiction.LiveBench—62.1%

Frequently asked questions

Is Qwen2.5-Coder (1.5B) better than Qwen3 8B?

Qwen3 8B has enough public results to be ranked (#238); Qwen2.5-Coder (1.5B) does not yet, so treat this comparison as directional.

How many benchmarks do Qwen2.5-Coder (1.5B) and Qwen3 8B share?

1 benchmark has published results for both models. Qwen2.5-Coder (1.5B) has 6 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper