Model comparison

ERNIE 5.1 vs GPT-5.4 nano

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 41.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Summary

  • They share 17 benchmarks with published results for both. ERNIE 5.1 scores higher in 5 categories and GPT-5.4 nano in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where ERNIE 5.1 leads 65.1 to 55.7.

Side by side

ERNIE 5.1 and GPT-5.4 nano specifications
ERNIE 5.1GPT-5.4 nano
ProviderBaiduOpenAI
Noometry Index43.841.9
Released—2026-03-17
WeightsProprietaryProprietary
Context window—400K
Max output—128K
Input $ / M tokens—$0.20
Output $ / M tokens—$1.25
Results tracked1940

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

ERNIE 5.1: 44.0 (#76), GPT-5.4 nano: 43.6 (#84)

Coding benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Coding14881405
SciCode—46.9%
WeirdML—49.2%
ALE-Bench—1,005

Agentic & Tool Use Not comparable

ERNIE 5.1: —, GPT-5.4 nano: —

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Search1227—

Reasoning GPT-5.4 nano leads

ERNIE 5.1: 21.9 (#211), GPT-5.4 nano: 23.7 (#173)

Reasoning benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Hard Prompts14811381
ARC-AGI-2—5.7%
Kagi LLM Benchmark—39.7%
NYT Connections (extended)23.4%—
ARC-AGI-1—51.5%
CritPt—9.3%
Chess Puzzles—30%
Mystery Game Puzzles—9%
DTBench—80.3%
LMCA—36.9%
Epoch Capabilities Index—145.81
ForecastBench—57.3

Math Too close to call

ERNIE 5.1: 40.3 (#92), GPT-5.4 nano: 40.9 (#88)

Math benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Math14811406
FrontierMath (Tiers 1-3)—44.9%
FrontierMath Tier 4—12.2%
OTIS Mock AIME 2024-2025—87.8%
ProofBench—5%
FrontierMath (Feb 2025 set)—25.9%
FrontierMath Tier 4 (v1)—6.3%

Knowledge Too close to call

ERNIE 5.1: 41.9 (#102), GPT-5.4 nano: 41.9 (#103)

Knowledge benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Expert14921396
GPQA Diamond—78.5%
SimpleQA Verified—11.7%
Vectara Hallucination Rate—3.1%

Multimodal Not comparable

ERNIE 5.1: —, GPT-5.4 nano: 36.7 (#78)

Multimodal benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Vision—1196

Multilingual ERNIE 5.1 leads

ERNIE 5.1: 55.5 (#29), GPT-5.4 nano: 48.6 (#140)

Multilingual benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Non-English14541359
LMArena Chinese15081392
LMArena French14881396
LMArena German14701367
LMArena Japanese14221343
LMArena Korean14271320
LMArena Russian14591363
LMArena Spanish14731371

Instruction Following ERNIE 5.1 leads

ERNIE 5.1: 76.7 (#37), GPT-5.4 nano: 71.9 (#144)

Instruction Following benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Instruction Following14601362

Long Context ERNIE 5.1 leads

ERNIE 5.1: 44.7 (#59), GPT-5.4 nano: 41.6 (#137)

Long Context benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Longer Query14621366

Writing & Preference ERNIE 5.1 leads

ERNIE 5.1: 65.1 (#52), GPT-5.4 nano: 55.7 (#142)

Writing & Preference benchmarks
BenchmarkERNIE 5.1GPT-5.4 nano
LMArena Text14681372
LMArena Creative Writing14411314
LMArena Multi-Turn14711382

Frequently asked questions

Is ERNIE 5.1 better than GPT-5.4 nano?

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 41.9 on the Noometry Index.

Is ERNIE 5.1 or GPT-5.4 nano better for coding?

They score almost the same on coding (44.0 vs 43.6); test both on your own repository before choosing.

How many benchmarks do ERNIE 5.1 and GPT-5.4 nano share?

17 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and GPT-5.4 nano has 40.

Related comparisons

Go deeper