Model comparison

Deepseek Coder v2 vs Qwen3.5 27B

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 35.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Deepseek Coder v2 DeepSeek

35.9

Rank #220 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Deepseek Coder v2 scores higher in 0 categories and Qwen3.5 27B in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 27B leads 59.3 to 38.2.

Side by side

Deepseek Coder v2 and Qwen3.5 27B specifications
Deepseek Coder v2Qwen3.5 27B
ProviderDeepSeekAlibaba (Qwen)
Noometry Index35.941.9
Released2024-06-172026-02-23
WeightsOpenOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.30
Output $ / M tokens—$2.40
Results tracked2428

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Deepseek Coder v2: 38.1 (#183), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Coding12511427
LMArena WebDev—1358
WeirdML—39.5%
BigCodeBench Instruct48.2%—
BigCodeBench Complete59.7%—
ALE-Bench—349.45
HumanEval+82.3%—
MBPP+75.1%—

Agentic & Tool Use Not comparable

Deepseek Coder v2: —, Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

Deepseek Coder v2: 23.6 (#176), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Hard Prompts12071414
NYT Connections (extended)—47.9%
Thematic Generalization—45.5%
DTBench—82.4%
LMCA—34%
WinoGrande83.7%—

Math Qwen3.5 27B leads

Deepseek Coder v2: 34.9 (#190), Qwen3.5 27B: 38.8 (#127)

Math benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Math12411429
MathArena Final-Answer Competitions—56.7%
GSM8K94.5%—

Knowledge Qwen3.5 27B leads

Deepseek Coder v2: 32.3 (#212), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Expert11811428
Vectara Hallucination Rate—12.1%
ARC (AI2) Challenge64.3%—

Multimodal Not comparable

Deepseek Coder v2: —, Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Vision—1241

Multilingual Qwen3.5 27B leads

Deepseek Coder v2: 36.3 (#240), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Non-English11821390
LMArena Chinese12011478
LMArena French11851410
LMArena German11641393
LMArena Japanese11261345
LMArena Korean11041358
LMArena Russian11881390
LMArena Spanish11531407

Instruction Following Qwen3.5 27B leads

Deepseek Coder v2: 61.7 (#242), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Instruction Following11801393

Long Context Qwen3.5 27B leads

Deepseek Coder v2: 37.0 (#224), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Longer Query12191413

Writing & Preference Qwen3.5 27B leads

Deepseek Coder v2: 38.2 (#253), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkDeepseek Coder v2Qwen3.5 27B
LMArena Text11911409
LMArena Creative Writing11201362
LMArena Multi-Turn11771410

Frequently asked questions

Is Deepseek Coder v2 better than Qwen3.5 27B?

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 35.9 on the Noometry Index.

Is Deepseek Coder v2 or Qwen3.5 27B better for coding?

They score almost the same on coding (38.1 vs 38.9); test both on your own repository before choosing.

How many benchmarks do Deepseek Coder v2 and Qwen3.5 27B share?

17 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper