Model comparison

DeepSeek V4.1 Flash vs ERNIE 5.1

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 43.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DeepSeek V4.1 Flash scores higher in 7 categories and ERNIE 5.1 in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4.1 Flash leads 50.2 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 89.6% for DeepSeek V4.1 Flash and 23.4% for ERNIE 5.1.
  • DeepSeek V4.1 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4.1 Flash and ERNIE 5.1 specifications
DeepSeek V4.1 FlashERNIE 5.1
ProviderDeepSeekBaidu
Noometry Index52.843.8
Released2026-09-09—
WeightsOpenProprietary
Context window1M—
Max output393K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked3719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 52.9 (#32), ERNIE 5.1: 44.0 (#76)

Coding benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Coding15061488
LMArena WebDev1619—
SciCode51.9%—
ALE-Bench1,092—

Agentic & Tool Use Not comparable

DeepSeek V4.1 Flash: 31.2 (#69), ERNIE 5.1: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
APEX-Agents39.5%—
GDP.pdf19.8%—
LMArena Search—1227

Reasoning DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 50.2 (#36), ERNIE 5.1: 21.9 (#211)

Reasoning benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
NYT Connections (extended)89.6%23.4%
LMArena Hard Prompts14831481
CritPt14.3%—
Mystery Game Puzzles43%—
DTBench89.9%—
LMCA47%—
Surface Evolver Bench46.3%—
Epoch Capabilities Index154.9—

Math DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 66.7 (#25), ERNIE 5.1: 40.3 (#92)

Math benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Math14771481
FrontierMath (Tiers 1-3)67.4%—
FrontierMath Tier 426.8%—
OTIS Mock AIME 2024-202598.3%—
ProofBench54%—

Knowledge DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 57.9 (#38), ERNIE 5.1: 41.9 (#102)

Knowledge benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Expert15061492
GPQA Diamond89.8%—

Multimodal Not comparable

DeepSeek V4.1 Flash: 39.1 (#61), ERNIE 5.1: —

Multimodal benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Vision1277—
Furniture Assembly34.2%—

Multilingual Too close to call

DeepSeek V4.1 Flash: 55.0 (#35), ERNIE 5.1: 55.5 (#29)

Multilingual benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Non-English14481454
LMArena Chinese14971508
LMArena French14521488
LMArena German14841470
LMArena Japanese14121422
LMArena Korean14521427
LMArena Russian14711459
LMArena Spanish14591473

Instruction Following Too close to call

DeepSeek V4.1 Flash: 77.3 (#26), ERNIE 5.1: 76.7 (#37)

Instruction Following benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Instruction Following14741460

Long Context Too close to call

DeepSeek V4.1 Flash: 45.2 (#47), ERNIE 5.1: 44.7 (#59)

Long Context benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Longer Query14751462

Writing & Preference Too close to call

DeepSeek V4.1 Flash: 65.4 (#48), ERNIE 5.1: 65.1 (#52)

Writing & Preference benchmarks
BenchmarkDeepSeek V4.1 FlashERNIE 5.1
LMArena Text14621468
LMArena Creative Writing14351441
LMArena Multi-Turn14571471
EQ-Bench Creative Writing1540—

Frequently asked questions

Is DeepSeek V4.1 Flash better than ERNIE 5.1?

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 43.8 on the Noometry Index.

Is DeepSeek V4.1 Flash or ERNIE 5.1 better for coding?

DeepSeek V4.1 Flash scores higher on coding benchmarks: 52.9 versus 44.0 in the Noometry coding category.

How many benchmarks do DeepSeek V4.1 Flash and ERNIE 5.1 share?

18 benchmarks have published results for both models. DeepSeek V4.1 Flash has 37 scored results on Noometry and ERNIE 5.1 has 19.

Related comparisons

Go deeper