Model comparison

ERNIE 5.1 vs GLM-4.5-Air

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 38.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Summary

  • They share 17 benchmarks with published results for both. ERNIE 5.1 scores higher in 7 categories and GLM-4.5-Air in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where ERNIE 5.1 leads 44.0 to 33.3.
  • GLM-4.5-Air has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and GLM-4.5-Air specifications
ERNIE 5.1GLM-4.5-Air
ProviderBaiduZ.ai (Zhipu)
Noometry Index43.838.9
Released—2025-07-20
WeightsProprietaryOpen
Context window—131K
Max output—98K
Input $ / M tokens—$0.20
Output $ / M tokens—$1.10
Results tracked1927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), GLM-4.5-Air: 33.3 (#259)

Coding benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Coding14881397
GSO—2.9%

Agentic & Tool Use Not comparable

ERNIE 5.1: —, GLM-4.5-Air: —

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Search1227—

Reasoning GLM-4.5-Air leads

ERNIE 5.1: 21.9 (#211), GLM-4.5-Air: 24.1 (#166)

Reasoning benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Hard Prompts14811379
Kagi LLM Benchmark—43%
NYT Connections (extended)23.4%—
ForecastBench—59.2

Math ERNIE 5.1 leads

ERNIE 5.1: 40.3 (#92), GLM-4.5-Air: 36.2 (#170)

Math benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Math14811396
Omni-MATH—39.1%

Knowledge ERNIE 5.1 leads

ERNIE 5.1: 41.9 (#102), GLM-4.5-Air: 35.0 (#191)

Knowledge benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Expert14921370
Humanity's Last Exam—8.1%
MMLU-Pro—76.2%
Vectara Hallucination Rate—9.3%
GPQA (HELM)—59.4%

Multilingual ERNIE 5.1 leads

ERNIE 5.1: 55.5 (#29), GLM-4.5-Air: 49.1 (#135)

Multilingual benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Non-English14541366
LMArena Chinese15081426
LMArena French14881399
LMArena German14701377
LMArena Japanese14221348
LMArena Korean14271308
LMArena Russian14591373
LMArena Spanish14731386

Instruction Following ERNIE 5.1 leads

ERNIE 5.1: 76.7 (#37), GLM-4.5-Air: 69.6 (#171)

Instruction Following benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Instruction Following14601354
IFEval—81.2%

Long Context ERNIE 5.1 leads

ERNIE 5.1: 44.7 (#59), GLM-4.5-Air: 41.6 (#135)

Long Context benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Longer Query14621366

Writing & Preference ERNIE 5.1 leads

ERNIE 5.1: 65.1 (#52), GLM-4.5-Air: 55.9 (#139)

Writing & Preference benchmarks
BenchmarkERNIE 5.1GLM-4.5-Air
LMArena Text14681384
LMArena Creative Writing14411343
LMArena Multi-Turn14711371
WildBench—78.9%

Frequently asked questions

Is ERNIE 5.1 better than GLM-4.5-Air?

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 38.9 on the Noometry Index.

Is ERNIE 5.1 or GLM-4.5-Air better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 33.3 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and GLM-4.5-Air share?

17 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and GLM-4.5-Air has 27.

Related comparisons

Go deeper