Model comparison

DeepSeek-V3.1-Terminus vs Qwen3.5 Plus

DeepSeek-V3.1-Terminus and Qwen3.5 Plus score almost the same on the Noometry Index (43.1 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • They share 3 benchmarks with published results for both. DeepSeek-V3.1-Terminus scores higher in 1 category and Qwen3.5 Plus in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.5 Plus leads 49.6 to 38.5.
  • The biggest single-benchmark swing is LMCA: 28.6% for DeepSeek-V3.1-Terminus and 36.4% for Qwen3.5 Plus.
  • DeepSeek-V3.1-Terminus is cheaper at $0.27 / $1 per million input/output tokens, against $0.40 / $2.40 for Qwen3.5 Plus.
  • Qwen3.5 Plus accepts more context: 1M tokens versus 164K.
  • DeepSeek-V3.1-Terminus has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1-Terminus and Qwen3.5 Plus specifications
DeepSeek-V3.1-TerminusQwen3.5 Plus
ProviderDeepSeekAlibaba (Qwen)
Noometry Index43.142.9
Released2025-09-222026-02-16
WeightsOpenProprietary
Context window164K1M
Max output147K66K
Input $ / M tokens$0.27$0.40
Output $ / M tokens$1$2.40
Results tracked1615

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek-V3.1-Terminus: 42.0 (#113), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
ALE-Bench745.17621.92
SciCode40.6%—
LMArena Coding1426—

Agentic & Tool Use Not comparable

DeepSeek-V3.1-Terminus: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

DeepSeek-V3.1-Terminus: 26.4 (#133), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
DTBench81.3%80.5%
LMCA28.6%36.4%
Kagi LLM Benchmark57.4%—
CritPt1.7%—
Chess Puzzles—22%
LMArena Hard Prompts1426—
Mystery Game Puzzles—17%
Epoch Capabilities Index—146.78

Math Qwen3.5 Plus leads

DeepSeek-V3.1-Terminus: 38.5 (#137), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
LMArena Math1402—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Not comparable

DeepSeek-V3.1-Terminus: —, Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
Vectara Hallucination Rate—10.7%

Multilingual Not comparable

DeepSeek-V3.1-Terminus: 52.1 (#92), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
LMArena Non-English1407—
LMArena Russian1436—

Instruction Following Not comparable

DeepSeek-V3.1-Terminus: 74.0 (#106), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
LMArena Instruction Following1404—

Long Context Too close to call

DeepSeek-V3.1-Terminus: 43.4 (#97), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1421—

Writing & Preference Not comparable

DeepSeek-V3.1-Terminus: 61.0 (#92), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusQwen3.5 Plus
LMArena Text1419—
LMArena Creative Writing1403—
LMArena Multi-Turn1411—

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than Qwen3.5 Plus?

DeepSeek-V3.1-Terminus and Qwen3.5 Plus score almost the same on the Noometry Index (43.1 vs 42.9), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.1-Terminus or Qwen3.5 Plus?

DeepSeek-V3.1-Terminus is cheaper. It lists at $0.27 per million input tokens and $1 per million output tokens; Qwen3.5 Plus lists at $0.40 and $2.40.

Which has the bigger context window?

Qwen3.5 Plus does, with 1M tokens against 164K.

How many benchmarks do DeepSeek-V3.1-Terminus and Qwen3.5 Plus share?

3 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper