Model comparison

DeepSeek LLM 67B vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 24.9 on the Noometry Index.

Last verified . 10 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and Qwen2.5 Plus 1127 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5 Plus 1127 leads 35.5 to 7.0.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and Qwen2.5 Plus 1127 specifications
DeepSeek LLM 67BQwen2.5 Plus 1127
ProviderDeepSeekAlibaba (Qwen)
Noometry Index24.938.8
Released2023-11-29—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1514

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 31.9 (#278), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Coding10961314

Reasoning Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 16.5 (#304), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Hard Prompts10701299
Chess Puzzles0%—
Epoch Capabilities Index110.5—

Math Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 8.7 (#324), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Math11081298
OTIS Mock AIME 2024-20250.8%—
MATH Level 56.4%—

Knowledge Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 7.0 (#313), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
GPQA Diamond24.6%—
LMArena Expert—1289

Multilingual Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 29.4 (#267), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Non-English10731265
LMArena Chinese11321314
LMArena German—1231
LMArena Japanese—1207
LMArena Russian—1271

Instruction Following Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 55.4 (#277), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Instruction Following10791275

Long Context Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 33.1 (#265), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Longer Query10921292

Writing & Preference Qwen2.5 Plus 1127 leads

DeepSeek LLM 67B: 31.6 (#282), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BQwen2.5 Plus 1127
LMArena Text11051299
LMArena Creative Writing10671262
LMArena Multi-Turn10821299

Frequently asked questions

Is DeepSeek LLM 67B better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Qwen2.5 Plus 1127 share?

10 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper