Model comparison

DeepSeek V4 Pro vs Qwen Plus

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 37.1 on the Noometry Index. Qwen Plus costs 1.6× less per token, which makes it the better buy when DeepSeek V4 Pro's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DeepSeek V4 Pro scores higher in 8 categories and Qwen Plus in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4 Pro leads 64.8 to 23.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 98.6% for DeepSeek V4 Pro and 17.8% for Qwen Plus.
  • Qwen Plus is cheaper at $0.40 / $1.20 per million input/output tokens, against $0.66 / $1.98 for DeepSeek V4 Pro.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4 Pro and Qwen Plus specifications
DeepSeek V4 ProQwen Plus
ProviderDeepSeekAlibaba (Qwen)
Noometry Index54.337.1
Released2026-04-242024-01-25
WeightsOpenProprietary
Context window1M1M
Max output393K33K
Input $ / M tokens$0.66$0.40
Output $ / M tokens$1.98$1.20
Results tracked4820

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

DeepSeek V4 Pro: 52.4 (#34), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
LMArena Coding14701328
SWE-bench Verified77.6%—
FrontierCode28.6%—
LMArena WebDev1582—
SciCode51%—
WeirdML66.2%—
ALE-Bench1,403—

Agentic & Tool Use Not comparable

DeepSeek V4 Pro: 32.8 (#58), Qwen Plus: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
APEX-Agents47.3%—
Vending-Bench 23,285—

Reasoning DeepSeek V4 Pro leads

DeepSeek V4 Pro: 56.5 (#24), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
Kagi LLM Benchmark53.5%63.3%
LMArena Hard Prompts14611317
DTBench93.9%81.1%
LMCA45.5%24%
ARC-AGI-261.3%—
NYT Connections (extended)91.3%—
ARC-AGI-190.5%—
CritPt18%—
Chess Puzzles47%—
Mystery Game Puzzles43%—
Surface Evolver Bench40%—
Epoch Capabilities Index155.31—
ForecastBench56.1—

Math DeepSeek V4 Pro leads

DeepSeek V4 Pro: 64.8 (#30), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
OTIS Mock AIME 2024-202598.6%17.8%
LMArena Math14551326
FrontierMath (Tiers 1-3)64.6%—
FrontierMath Tier 426.8%—
MathArena Final-Answer Competitions76.6%—
ProofBench50%—
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%

Knowledge DeepSeek V4 Pro leads

DeepSeek V4 Pro: 59.5 (#31), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
GPQA Diamond91.7%48.1%
LMArena Expert14641328
SimpleQA Verified52.9%—
Vectara Hallucination Rate8.6%—

Multilingual DeepSeek V4 Pro leads

DeepSeek V4 Pro: 54.4 (#45), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
LMArena Non-English14391310
LMArena Chinese14861347
LMArena Japanese14451251
LMArena Russian14531323
LMArena French1472—
LMArena German1458—
LMArena Korean1447—
LMArena Spanish1458—

Instruction Following DeepSeek V4 Pro leads

DeepSeek V4 Pro: 76.1 (#47), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
LMArena Instruction Following14481303

Long Context DeepSeek V4 Pro leads

DeepSeek V4 Pro: 45.0 (#51), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
LMArena Longer Query14581324
CL-bench Life13.5%—

Writing & Preference DeepSeek V4 Pro leads

DeepSeek V4 Pro: 65.5 (#46), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 ProQwen Plus
LMArena Text14511326
LMArena Creative Writing14461293
LMArena Multi-Turn14671336
EQ-Bench Creative Writing1553—
EQ-Bench 41166—

Frequently asked questions

Is DeepSeek V4 Pro better than Qwen Plus?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 37.1 on the Noometry Index. Qwen Plus costs 1.6× less per token, which makes it the better buy when DeepSeek V4 Pro's lead doesn't matter for your workload.

Which is cheaper, DeepSeek V4 Pro or Qwen Plus?

Qwen Plus is cheaper. It lists at $0.40 per million input tokens and $1.20 per million output tokens; DeepSeek V4 Pro lists at $0.66 and $1.98.

Is DeepSeek V4 Pro or Qwen Plus better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 38.9 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do DeepSeek V4 Pro and Qwen Plus share?

18 benchmarks have published results for both models. DeepSeek V4 Pro has 48 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper