Model comparison

DeepSeek-R1 vs Qwen3.5 27B

DeepSeek-R1 and Qwen3.5 27B score almost the same on the Noometry Index (42.3 vs 41.9), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 20 benchmarks with published results for both. DeepSeek-R1 scores higher in 6 categories and Qwen3.5 27B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5 27B leads 27.5 to 18.6.
  • Qwen3.5 27B is cheaper at $0.30 / $2.40 per million input/output tokens, against $0.50 / $2.15 for DeepSeek-R1.
  • Qwen3.5 27B accepts more context: 262K tokens versus 164K.
  • Qwen3.5 27B has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1 and Qwen3.5 27B specifications
DeepSeek-R1Qwen3.5 27B
ProviderDeepSeekAlibaba (Qwen)
Noometry Index42.341.9
Released2025-01-202026-02-23
WeightsProprietaryOpen
Context window164K262K
Max output64K66K
Input $ / M tokens$0.50$0.30
Output $ / M tokens$2.15$2.40
Results tracked5228

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

DeepSeek-R1: 46.3 (#68), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
WeirdML41.6%39.5%
LMArena Coding14271427
ALE-Bench804.12349.45
Aider Polyglot71.4%—
LMArena WebDev—1358
SciCode35.7%—
LiveBench Coding66.7%—
AlgoTune1.7—

Agentic & Tool Use Not comparable

DeepSeek-R1: 30.7 (#75), Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
DeepResearch Bench35.1%—
BALROG34.9%—
METR Time Horizons53.8%—
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

DeepSeek-R1: 18.6 (#278), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Hard Prompts14161414
ARC-AGI-21.3%—
SimpleBench40.8%—
Kagi LLM Benchmark69.4%—
NYT Connections (extended)—47.9%
ARC-AGI-121.2%—
CritPt1.1%—
Thematic Generalization—45.5%
LiveBench Reasoning83.2%—
DTBench—82.4%
LiveBench Data Analysis69.8%—
LMCA—34%
Epoch Capabilities Index141.29—
ForecastBench60—
LiveBench71.6%—

Math DeepSeek-R1 leads

DeepSeek-R1: 43.8 (#79), Qwen3.5 27B: 38.8 (#127)

Math benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Math14001429
MathArena Final-Answer Competitions—56.7%
OTIS Mock AIME 2024-202566.4%—
Omni-MATH42.4%—
LiveBench Math80.7%—
MATH Level 596.6%—

Knowledge DeepSeek-R1 leads

DeepSeek-R1: 44.5 (#87), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
Vectara Hallucination Rate11.3%12.1%
LMArena Expert13941428
GPQA Diamond76.3%—
MMLU-Pro79.3%—
Confabulations12.7%—
GPQA (HELM)66.6%—

Multimodal Not comparable

DeepSeek-R1: —, Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Vision—1241

Multilingual DeepSeek-R1 leads

DeepSeek-R1: 52.4 (#85), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Non-English14121390
LMArena Chinese14421478
LMArena French14171410
LMArena German14041393
LMArena Japanese13911345
LMArena Korean13601358
LMArena Russian14231390
LMArena Spanish14111407

Instruction Following Qwen3.5 27B leads

DeepSeek-R1: 72.0 (#143), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Instruction Following13821393
LiveBench Instruction Following80.5%—
IFEval78.4%—

Long Context DeepSeek-R1 leads

DeepSeek-R1: 45.4 (#36), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Longer Query13911413
Fiction.LiveBench75%—

Writing & Preference DeepSeek-R1 leads

DeepSeek-R1: 61.4 (#88), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1Qwen3.5 27B
LMArena Text14281409
LMArena Creative Writing14051362
LMArena Multi-Turn14051410
Short-Story Creative Writing83%—
EQ-Bench Creative Writing1500—
WildBench82.8%—
LiveBench Language48.5%—

Frequently asked questions

Is DeepSeek-R1 better than Qwen3.5 27B?

DeepSeek-R1 and Qwen3.5 27B score almost the same on the Noometry Index (42.3 vs 41.9), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-R1 or Qwen3.5 27B?

Qwen3.5 27B is cheaper. It lists at $0.30 per million input tokens and $2.40 per million output tokens; DeepSeek-R1 lists at $0.50 and $2.15.

Is DeepSeek-R1 or Qwen3.5 27B better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 38.9 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5 27B does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-R1 and Qwen3.5 27B share?

20 benchmarks have published results for both models. DeepSeek-R1 has 52 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper