Model comparison

DeepSeek-V3.1-Terminus vs ERNIE 5.0 0110

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 41.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

ERNIE 5.0 0110 Baidu

41.8

Rank #129 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DeepSeek-V3.1-Terminus scores higher in 1 category and ERNIE 5.0 0110 in 6 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.1-Terminus leads 26.4 to 17.0.
  • DeepSeek-V3.1-Terminus has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1-Terminus and ERNIE 5.0 0110 specifications
DeepSeek-V3.1-TerminusERNIE 5.0 0110
ProviderDeepSeekBaidu
Noometry Index43.141.8
Released2025-09-22—
WeightsOpenProprietary
Context window164K—
Max output147K—
Input $ / M tokens$0.27—
Output $ / M tokens$1—
Results tracked1620

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V3.1-Terminus: 42.0 (#113), ERNIE 5.0 0110: 43.0 (#94)

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Coding14261455
SciCode40.6%—
ALE-Bench745.17—

Reasoning DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 26.4 (#133), ERNIE 5.0 0110: 17.0 (#297)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Hard Prompts14261445
Kagi LLM Benchmark57.4%—
NYT Connections (extended)—10.3%
CritPt1.7%—
Thematic Generalization—41.7%
DTBench81.3%—
LMCA28.6%—

Math Too close to call

DeepSeek-V3.1-Terminus: 38.5 (#137), ERNIE 5.0 0110: 39.3 (#110)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Math14021437

Knowledge Not comparable

DeepSeek-V3.1-Terminus: —, ERNIE 5.0 0110: 39.8 (#128)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Expert—1428

Multimodal Not comparable

DeepSeek-V3.1-Terminus: —, ERNIE 5.0 0110: 39.9 (#53)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Vision—1249

Multilingual ERNIE 5.0 0110 leads

DeepSeek-V3.1-Terminus: 52.1 (#92), ERNIE 5.0 0110: 54.1 (#49)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Non-English14071436
LMArena Russian14361446
LMArena Chinese—1512
LMArena French—1467
LMArena German—1460
LMArena Japanese—1382
LMArena Korean—1406
LMArena Spanish—1473

Instruction Following Too close to call

DeepSeek-V3.1-Terminus: 74.0 (#106), ERNIE 5.0 0110: 74.5 (#92)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Instruction Following14041413

Long Context Too close to call

DeepSeek-V3.1-Terminus: 43.4 (#97), ERNIE 5.0 0110: 43.4 (#95)

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Longer Query14211422

Writing & Preference ERNIE 5.0 0110 leads

DeepSeek-V3.1-Terminus: 61.0 (#92), ERNIE 5.0 0110: 63.1 (#66)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusERNIE 5.0 0110
LMArena Text14191445
LMArena Creative Writing14031426
LMArena Multi-Turn14111434

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than ERNIE 5.0 0110?

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 41.8 on the Noometry Index.

Is DeepSeek-V3.1-Terminus or ERNIE 5.0 0110 better for coding?

They score almost the same on coding (42.0 vs 43.0); test both on your own repository before choosing.

How many benchmarks do DeepSeek-V3.1-Terminus and ERNIE 5.0 0110 share?

10 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and ERNIE 5.0 0110 has 20.

Related comparisons

Go deeper