Model comparison

ERNIE 5.0 0110 vs GPT-4

ERNIE 5.0 0110 is the stronger model overall, scoring 41.8 to 29.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

ERNIE 5.0 0110 Baidu

41.8

Rank #129 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 17 benchmarks with published results for both. ERNIE 5.0 0110 scores higher in 7 categories and GPT-4 in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where ERNIE 5.0 0110 leads 39.3 to 10.8.

Side by side

ERNIE 5.0 0110 and GPT-4 specifications
ERNIE 5.0 0110GPT-4
ProviderBaiduOpenAI
Noometry Index41.829.1
Released—2023-03-14
WeightsProprietaryProprietary
Context window—8K
Max output—8K
Input $ / M tokens—$30
Output $ / M tokens—$60
Results tracked2038

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 43.0 (#94), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Coding14551254
WeirdML—12.4%
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
HumanEval+—79.3%

Agentic & Tool Use Not comparable

ERNIE 5.0 0110: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.0 0110GPT-4
METR Time Horizons—36.1%

Reasoning Too close to call

ERNIE 5.0 0110: 17.0 (#297), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Hard Prompts14451241
NYT Connections (extended)10.3%—
Chess Puzzles—4%
Thematic Generalization41.7%—
Mystery Game Puzzles—12%
DTBench—62.7%
LMCA—17.1%
BIG-Bench Hard—75.1%
Epoch Capabilities Index—125.89
ForecastBench—57.8
HellaSwag—95.3%
WinoGrande—87.5%

Math ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 39.3 (#110), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Math14371269
OTIS Mock AIME 2024-2025—1.1%
MATH Level 5—23%
GSM8K—92%

Knowledge ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 39.8 (#128), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Expert14281211
GPQA Diamond—35.7%
MMLU—86.4%
TriviaQA—84.8%

Multimodal Not comparable

ERNIE 5.0 0110: 39.9 (#53), GPT-4: —

Multimodal benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Vision1249—

Multilingual ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 54.1 (#49), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Non-English14361246
LMArena Chinese15121242
LMArena French14671283
LMArena German14601251
LMArena Japanese13821209
LMArena Korean14061184
LMArena Russian14461251
LMArena Spanish14731261

Instruction Following ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 74.5 (#92), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Instruction Following14131241

Long Context ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 43.4 (#95), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Longer Query14221244

Writing & Preference ERNIE 5.0 0110 leads

ERNIE 5.0 0110: 63.1 (#66), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkERNIE 5.0 0110GPT-4
LMArena Text14451263
LMArena Creative Writing14261244
LMArena Multi-Turn14341257
EQ-Bench Creative Writing—752

Frequently asked questions

Is ERNIE 5.0 0110 better than GPT-4?

ERNIE 5.0 0110 is the stronger model overall, scoring 41.8 to 29.1 on the Noometry Index.

Is ERNIE 5.0 0110 or GPT-4 better for coding?

ERNIE 5.0 0110 scores higher on coding benchmarks: 43.0 versus 31.6 in the Noometry coding category.

How many benchmarks do ERNIE 5.0 0110 and GPT-4 share?

17 benchmarks have published results for both models. ERNIE 5.0 0110 has 20 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper