Model comparison

DeepSeek-V3.1-Terminus vs Hy3

Hy3 is the stronger model overall, scoring 44.2 to 43.1 on the Noometry Index.

Last verified . 10 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DeepSeek-V3.1-Terminus scores higher in 1 category and Hy3 in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Hy3 leads 46.8 to 42.0.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $0.27 / $1 for DeepSeek-V3.1-Terminus.
  • Hy3 accepts more context: 262K tokens versus 164K.

Side by side

DeepSeek-V3.1-Terminus and Hy3 specifications
DeepSeek-V3.1-TerminusHy3
ProviderDeepSeekTencent
Noometry Index43.144.2
Released2025-09-222026-07-06
WeightsOpenOpen
Context window164K262K
Max output147K128K
Input $ / M tokens$0.27$0.0825
Output $ / M tokens$1$0.33
Results tracked1619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

DeepSeek-V3.1-Terminus: 42.0 (#113), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Coding14261464
LMArena WebDev—1508
SciCode40.6%—
ALE-Bench745.17—

Reasoning Too close to call

DeepSeek-V3.1-Terminus: 26.4 (#133), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Hard Prompts14261447
Kagi LLM Benchmark57.4%—
NYT Connections (extended)—41.2%
CritPt1.7%—
DTBench81.3%—
LMCA28.6%—

Math Hy3 leads

DeepSeek-V3.1-Terminus: 38.5 (#137), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Math14021475

Knowledge Not comparable

DeepSeek-V3.1-Terminus: —, Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Expert—1460

Multilingual Hy3 leads

DeepSeek-V3.1-Terminus: 52.1 (#92), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Non-English14071426
LMArena Russian14361432
LMArena Chinese—1493
LMArena French—1461
LMArena German—1439
LMArena Japanese—1392
LMArena Korean—1395
LMArena Spanish—1456

Instruction Following Hy3 leads

DeepSeek-V3.1-Terminus: 74.0 (#106), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Instruction Following14041426

Long Context Too close to call

DeepSeek-V3.1-Terminus: 43.4 (#97), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Longer Query14211442

Writing & Preference Hy3 leads

DeepSeek-V3.1-Terminus: 61.0 (#92), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusHy3
LMArena Text14191439
LMArena Creative Writing14031402
LMArena Multi-Turn14111436

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than Hy3?

Hy3 is the stronger model overall, scoring 44.2 to 43.1 on the Noometry Index.

Which is cheaper, DeepSeek-V3.1-Terminus or Hy3?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; DeepSeek-V3.1-Terminus lists at $0.27 and $1.

Is DeepSeek-V3.1-Terminus or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 42.0 in the Noometry coding category.

Which has the bigger context window?

Hy3 does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-V3.1-Terminus and Hy3 share?

10 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper