Model comparison

DeepSeek-V3.1 vs ERNIE 5.0 0110

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 41.8 on the Noometry Index.

Last verified . 17 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

ERNIE 5.0 0110 Baidu

41.8

Rank #129 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 2 categories and ERNIE 5.0 0110 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.1 leads 27.9 to 17.0.
  • DeepSeek-V3.1 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1 and ERNIE 5.0 0110 specifications
DeepSeek-V3.1ERNIE 5.0 0110
ProviderDeepSeekBaidu
Noometry Index42.841.8
Released2025-08-21—
WeightsOpenProprietary
Context window164K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.95—
Results tracked2720

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.0 0110 leads

DeepSeek-V3.1: 40.3 (#144), ERNIE 5.0 0110: 43.0 (#94)

Coding benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Coding14171455
WeirdML38.4%—

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), ERNIE 5.0 0110: 17.0 (#297)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Hard Prompts14171445
SimpleBench40%—
Kagi LLM Benchmark53.2%—
NYT Connections (extended)—10.3%
Thematic Generalization—41.7%
DTBench82.7%—
LMCA24.3%—
Epoch Capabilities Index139.92—
ForecastBench58—

Math Too close to call

DeepSeek-V3.1: 38.9 (#122), ERNIE 5.0 0110: 39.3 (#110)

Math benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Math14201437

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), ERNIE 5.0 0110: 39.8 (#128)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Expert14051428
Vectara Hallucination Rate5.5%—

Multimodal Not comparable

DeepSeek-V3.1: —, ERNIE 5.0 0110: 39.9 (#53)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Vision—1249

Multilingual ERNIE 5.0 0110 leads

DeepSeek-V3.1: 51.6 (#106), ERNIE 5.0 0110: 54.1 (#49)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Non-English14001436
LMArena Chinese14691512
LMArena French14471467
LMArena German14111460
LMArena Japanese13781382
LMArena Korean13371406
LMArena Russian14051446
LMArena Spanish14311473

Instruction Following Too close to call

DeepSeek-V3.1: 73.9 (#110), ERNIE 5.0 0110: 74.5 (#92)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Instruction Following14001413

Long Context ERNIE 5.0 0110 leads

DeepSeek-V3.1: 36.3 (#232), ERNIE 5.0 0110: 43.4 (#95)

Long Context benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Longer Query14221422
Fiction.LiveBench52.8%—

Writing & Preference ERNIE 5.0 0110 leads

DeepSeek-V3.1: 60.3 (#98), ERNIE 5.0 0110: 63.1 (#66)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1ERNIE 5.0 0110
LMArena Text14201445
LMArena Creative Writing14011426
LMArena Multi-Turn14081434
EQ-Bench Creative Writing1436—

Frequently asked questions

Is DeepSeek-V3.1 better than ERNIE 5.0 0110?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 41.8 on the Noometry Index.

Is DeepSeek-V3.1 or ERNIE 5.0 0110 better for coding?

ERNIE 5.0 0110 scores higher on coding benchmarks: 43.0 versus 40.3 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.1 and ERNIE 5.0 0110 share?

17 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and ERNIE 5.0 0110 has 20.

Related comparisons

Go deeper