Model comparison

ERNIE 5.1 vs Kimi K2.6

Kimi K2.6 is the stronger model overall, scoring 47.7 to 43.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 18 benchmarks with published results for both. ERNIE 5.1 scores higher in 2 categories and Kimi K2.6 in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K2.6 leads 40.5 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 87.2% for Kimi K2.6.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Kimi K2.6 specifications
ERNIE 5.1Kimi K2.6
ProviderBaiduMoonshot AI
Noometry Index43.847.7
Released—2026-04-20
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.95
Output $ / M tokens—$4
Results tracked1951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.6 leads

ERNIE 5.1: 44.0 (#76), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Coding14881488
SWE-bench Verified—76.7%
LMArena WebDev—1509
SciCode—53.5%
WeirdML—55.9%
ALE-Bench—1,093

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Kimi K2.6
OSWorld 2.0—4.6%
ExploitBench—18.4%
GBAEval—0.9%
GDP.pdf—12%
LMArena Search1227—
Vending-Bench 2—6,205

Reasoning Kimi K2.6 leads

ERNIE 5.1: 21.9 (#211), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkERNIE 5.1Kimi K2.6
NYT Connections (extended)23.4%87.2%
LMArena Hard Prompts14811470
CritPt—8%
Chess Puzzles—26%
EBR-Bench—2.4%
Mystery Game Puzzles—18%
DTBench—90.9%
LMCA—37.3%
Epoch Capabilities Index—151.05

Math Kimi K2.6 leads

ERNIE 5.1: 40.3 (#92), Kimi K2.6: 57.0 (#41)

Knowledge Kimi K2.6 leads

ERNIE 5.1: 41.9 (#102), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Expert14921491
GPQA Diamond—90.8%
SimpleQA Verified—34.9%
Vectara Hallucination Rate—10.8%

Multimodal Not comparable

ERNIE 5.1: —, Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Vision—1283
Blueprint-Bench 2—3.9%
Furniture Assembly—21.7%
LMArena Document—1451

Multilingual Too close to call

ERNIE 5.1: 55.5 (#29), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Non-English14541446
LMArena Chinese15081521
LMArena French14881471
LMArena German14701450
LMArena Japanese14221443
LMArena Korean14271427
LMArena Russian14591446
LMArena Spanish14731464

Instruction Following Too close to call

ERNIE 5.1: 76.7 (#37), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Instruction Following14601451

Long Context Too close to call

ERNIE 5.1: 44.7 (#59), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Longer Query14621468

Writing & Preference Kimi K2.6 leads

ERNIE 5.1: 65.1 (#52), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Kimi K2.6
LMArena Text14681455
LMArena Creative Writing14411434
LMArena Multi-Turn14711453
EQ-Bench Creative Writing—1725
EQ-Bench 4—1202

Frequently asked questions

Is ERNIE 5.1 better than Kimi K2.6?

Kimi K2.6 is the stronger model overall, scoring 47.7 to 43.8 on the Noometry Index.

Is ERNIE 5.1 or Kimi K2.6 better for coding?

Kimi K2.6 scores higher on coding benchmarks: 50.7 versus 44.0 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Kimi K2.6 share?

18 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper