Model comparison

DeepSeek-V3.1 vs Qwen3.5-9B

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 33.8 on the Noometry Index. Qwen3.5-9B costs 3.8× less per token, which makes it the better buy when DeepSeek-V3.1's lead doesn't matter for your workload.

Last verified . 3 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Qwen3.5-9B Alibaba (Qwen)

33.8

Rank #236 Confirmed

Summary

  • They share 3 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 3 categories and Qwen3.5-9B in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.1 leads 27.9 to 23.1.
  • The biggest single-benchmark swing is DTBench: 82.7% for DeepSeek-V3.1 and 71.2% for Qwen3.5-9B.
  • Qwen3.5-9B is cheaper at $0.10 / $0.15 per million input/output tokens, against $0.25 / $0.95 for DeepSeek-V3.1.
  • Qwen3.5-9B accepts more context: 262K tokens versus 164K.

Side by side

DeepSeek-V3.1 and Qwen3.5-9B specifications
DeepSeek-V3.1Qwen3.5-9B
ProviderDeepSeekAlibaba (Qwen)
Noometry Index42.833.8
Released2025-08-212026-02-23
WeightsOpenOpen
Context window164K262K
Max output8K66K
Input $ / M tokens$0.25$0.10
Output $ / M tokens$0.95$0.15
Results tracked2710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1 leads

DeepSeek-V3.1: 40.3 (#144), Qwen3.5-9B: 35.9 (#217)

Coding benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
SciCode—27.5%
WeirdML38.4%—
LMArena Coding1417—

Agentic & Tool Use Not comparable

DeepSeek-V3.1: —, Qwen3.5-9B: 14.5 (#151)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
Terminal-Bench—9.2%

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), Qwen3.5-9B: 23.1 (#182)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
DTBench82.7%71.2%
LMCA24.3%24.5%
Epoch Capabilities Index139.92139.46
SimpleBench40%—
Kagi LLM Benchmark53.2%—
CritPt—0.3%
Chess Puzzles—12%
LMArena Hard Prompts1417—
ForecastBench58—

Math DeepSeek-V3.1 leads

DeepSeek-V3.1: 38.9 (#122), Qwen3.5-9B: 34.8 (#192)

Math benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
MathArena Final-Answer Competitions—48.5%
OTIS Mock AIME 2024-2025—61.7%
LMArena Math1420—

Knowledge Qwen3.5-9B leads

DeepSeek-V3.1: 43.7 (#90), Qwen3.5-9B: 46.0 (#84)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
GPQA Diamond—79%
Vectara Hallucination Rate5.5%—
LMArena Expert1405—

Multilingual Not comparable

DeepSeek-V3.1: 51.6 (#106), Qwen3.5-9B: —

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
LMArena Non-English1400—
LMArena Chinese1469—
LMArena French1447—
LMArena German1411—
LMArena Japanese1378—
LMArena Korean1337—
LMArena Russian1405—
LMArena Spanish1431—

Instruction Following Not comparable

DeepSeek-V3.1: 73.9 (#110), Qwen3.5-9B: —

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
LMArena Instruction Following1400—

Long Context Not comparable

DeepSeek-V3.1: 36.3 (#232), Qwen3.5-9B: —

Long Context benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
Fiction.LiveBench52.8%—
LMArena Longer Query1422—

Writing & Preference Not comparable

DeepSeek-V3.1: 60.3 (#98), Qwen3.5-9B: —

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5-9B
LMArena Text1420—
LMArena Creative Writing1401—
EQ-Bench Creative Writing1436—
LMArena Multi-Turn1408—

Frequently asked questions

Is DeepSeek-V3.1 better than Qwen3.5-9B?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 33.8 on the Noometry Index. Qwen3.5-9B costs 3.8× less per token, which makes it the better buy when DeepSeek-V3.1's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3.1 or Qwen3.5-9B?

Qwen3.5-9B is cheaper. It lists at $0.10 per million input tokens and $0.15 per million output tokens; DeepSeek-V3.1 lists at $0.25 and $0.95.

Is DeepSeek-V3.1 or Qwen3.5-9B better for coding?

DeepSeek-V3.1 scores higher on coding benchmarks: 40.3 versus 35.9 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5-9B does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-V3.1 and Qwen3.5-9B share?

3 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Qwen3.5-9B has 10.

Related comparisons

Go deeper