Model comparison

DBRX vs GLM-4.5-Air

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 29.4 on the Noometry Index.

Last verified . 17 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DBRX scores higher in 0 categories and GLM-4.5-Air in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GLM-4.5-Air leads 55.9 to 33.6.

Side by side

DBRX and GLM-4.5-Air specifications
DBRXGLM-4.5-Air
ProviderDatabricksZ.ai (Zhipu)
Noometry Index29.438.9
Released2024-03-272025-07-20
WeightsOpenOpen
Context window—131K
Max output—98K
Input $ / M tokens—$0.20
Output $ / M tokens—$1.10
Results tracked2127

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DBRX: 32.9 (#266), GLM-4.5-Air: 33.3 (#259)

Coding benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Coding11321397
GSO—2.9%
HumanEval+70.1%—
MBPP+55.8%—

Reasoning GLM-4.5-Air leads

DBRX: 21.4 (#222), GLM-4.5-Air: 24.1 (#166)

Reasoning benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Hard Prompts11131379
Kagi LLM Benchmark—43%
ForecastBench—59.2

Math GLM-4.5-Air leads

DBRX: 24.3 (#269), GLM-4.5-Air: 36.2 (#170)

Math benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Math11451396
Omni-MATH—39.1%
MATH Level 511.7%—

Knowledge GLM-4.5-Air leads

DBRX: 14.9 (#294), GLM-4.5-Air: 35.0 (#191)

Knowledge benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Expert10761370
GPQA Diamond32.9%—
Humanity's Last Exam—8.1%
MMLU-Pro—76.2%
Vectara Hallucination Rate—9.3%
GPQA (HELM)—59.4%

Multilingual GLM-4.5-Air leads

DBRX: 29.3 (#268), GLM-4.5-Air: 49.1 (#135)

Multilingual benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Non-English10711366
LMArena Chinese10681426
LMArena French10961399
LMArena German10571377
LMArena Japanese9901348
LMArena Korean9931308
LMArena Russian10781373
LMArena Spanish10641386

Instruction Following GLM-4.5-Air leads

DBRX: 57.5 (#270), GLM-4.5-Air: 69.6 (#171)

Instruction Following benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Instruction Following11121354
IFEval—81.2%

Long Context GLM-4.5-Air leads

DBRX: 33.7 (#258), GLM-4.5-Air: 41.6 (#135)

Long Context benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Longer Query11121366

Writing & Preference GLM-4.5-Air leads

DBRX: 33.6 (#275), GLM-4.5-Air: 55.9 (#139)

Writing & Preference benchmarks
BenchmarkDBRXGLM-4.5-Air
LMArena Text11191384
LMArena Creative Writing11041343
LMArena Multi-Turn11111371
WildBench—78.9%

Frequently asked questions

Is DBRX better than GLM-4.5-Air?

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 29.4 on the Noometry Index.

Is DBRX or GLM-4.5-Air better for coding?

They score almost the same on coding (32.9 vs 33.3); test both on your own repository before choosing.

How many benchmarks do DBRX and GLM-4.5-Air share?

17 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and GLM-4.5-Air has 27.

Related comparisons

Go deeper