Model comparison

Qwen2.5 Plus 1127 vs Qwen3 14B

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 35.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Qwen3 14B Alibaba (Qwen)

35.5

Rank #225 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen2.5 Plus 1127 leads 25.9 to 18.5.
  • Qwen3 14B has downloadable open weights; the other is API-only.

Side by side

Qwen2.5 Plus 1127 and Qwen3 14B specifications
Qwen2.5 Plus 1127Qwen3 14B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index38.835.5
Released—2025-04
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.35
Output $ / M tokens—$1.40
Results tracked1412

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 38.5 (#175), Qwen3 14B: 37.3 (#195)

Coding benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
SciCode—31.6%
LMArena Coding1314—

Agentic & Tool Use Not comparable

Qwen2.5 Plus 1127: —, Qwen3 14B: 29.6 (#83)

Agentic & Tool Use benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
Berkeley Function Calling Leaderboard—41%

Reasoning Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 25.9 (#141), Qwen3 14B: 18.5 (#280)

Reasoning benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
Kagi LLM Benchmark—49.1%
CritPt—0%
Chess Puzzles—4%
LMArena Hard Prompts1299—
DTBench—64%
LMCA—18.2%
Epoch Capabilities Index—138.23

Math Qwen3 14B leads

Qwen2.5 Plus 1127: 36.1 (#174), Qwen3 14B: 38.6 (#133)

Math benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
OTIS Mock AIME 2024-2025—66.4%
LMArena Math1298—

Knowledge Qwen3 14B leads

Qwen2.5 Plus 1127: 35.5 (#183), Qwen3 14B: 39.3 (#134)

Knowledge benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
GPQA Diamond—63.8%
Vectara Hallucination Rate—5.4%
LMArena Expert1289—

Multilingual Not comparable

Qwen2.5 Plus 1127: 41.9 (#201), Qwen3 14B: —

Multilingual benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
LMArena Non-English1265—
LMArena Chinese1314—
LMArena German1231—
LMArena Japanese1207—
LMArena Russian1271—

Instruction Following Not comparable

Qwen2.5 Plus 1127: 67.2 (#199), Qwen3 14B: —

Instruction Following benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
LMArena Instruction Following1275—

Long Context Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 39.2 (#184), Qwen3 14B: 38.1 (#204)

Long Context benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
Fiction.LiveBench—62.5%
LMArena Longer Query1292—

Writing & Preference Not comparable

Qwen2.5 Plus 1127: 49.4 (#192), Qwen3 14B: —

Writing & Preference benchmarks
BenchmarkQwen2.5 Plus 1127Qwen3 14B
LMArena Text1299—
LMArena Creative Writing1262—
LMArena Multi-Turn1299—

Frequently asked questions

Is Qwen2.5 Plus 1127 better than Qwen3 14B?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 35.5 on the Noometry Index.

Is Qwen2.5 Plus 1127 or Qwen3 14B better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 37.3 in the Noometry coding category.

How many benchmarks do Qwen2.5 Plus 1127 and Qwen3 14B share?

0 benchmarks have published results for both models. Qwen2.5 Plus 1127 has 14 scored results on Noometry and Qwen3 14B has 12.

Related comparisons

Go deeper