Model comparison

GLM-4.5-Air vs GPT-4

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 29.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 18 benchmarks with published results for both. GLM-4.5-Air scores higher in 8 categories and GPT-4 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-4.5-Air leads 36.2 to 10.8.
  • GLM-4.5-Air is cheaper at $0.20 / $1.10 per million input/output tokens, against $30 / $60 for GPT-4.
  • GLM-4.5-Air accepts more context: 131K tokens versus 8K.
  • GLM-4.5-Air has downloadable open weights; the other is API-only.

Side by side

GLM-4.5-Air and GPT-4 specifications
GLM-4.5-AirGPT-4
ProviderZ.ai (Zhipu)OpenAI
Noometry Index38.929.1
Released2025-07-202023-03-14
WeightsOpenProprietary
Context window131K8K
Max output98K8K
Input $ / M tokens$0.20$30
Output $ / M tokens$1.10$60
Results tracked2738

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.5-Air leads

GLM-4.5-Air: 33.3 (#259), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Coding13971254
GSO2.9%—
WeirdML—12.4%
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
HumanEval+—79.3%

Agentic & Tool Use Not comparable

GLM-4.5-Air: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkGLM-4.5-AirGPT-4
METR Time Horizons—36.1%

Reasoning GLM-4.5-Air leads

GLM-4.5-Air: 24.1 (#166), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Hard Prompts13791241
ForecastBench59.257.8
Kagi LLM Benchmark43%—
Chess Puzzles—4%
Mystery Game Puzzles—12%
DTBench—62.7%
LMCA—17.1%
BIG-Bench Hard—75.1%
Epoch Capabilities Index—125.89
HellaSwag—95.3%
WinoGrande—87.5%

Math GLM-4.5-Air leads

GLM-4.5-Air: 36.2 (#170), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Math13961269
OTIS Mock AIME 2024-2025—1.1%
Omni-MATH39.1%—
MATH Level 5—23%
GSM8K—92%

Knowledge GLM-4.5-Air leads

GLM-4.5-Air: 35.0 (#191), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Expert13701211
GPQA Diamond—35.7%
Humanity's Last Exam8.1%—
MMLU-Pro76.2%—
Vectara Hallucination Rate9.3%—
GPQA (HELM)59.4%—
MMLU—86.4%
TriviaQA—84.8%

Multilingual GLM-4.5-Air leads

GLM-4.5-Air: 49.1 (#135), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Non-English13661246
LMArena Chinese14261242
LMArena French13991283
LMArena German13771251
LMArena Japanese13481209
LMArena Korean13081184
LMArena Russian13731251
LMArena Spanish13861261

Instruction Following GLM-4.5-Air leads

GLM-4.5-Air: 69.6 (#171), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Instruction Following13541241
IFEval81.2%—

Long Context GLM-4.5-Air leads

GLM-4.5-Air: 41.6 (#135), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Longer Query13661244

Writing & Preference GLM-4.5-Air leads

GLM-4.5-Air: 55.9 (#139), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirGPT-4
LMArena Text13841263
LMArena Creative Writing13431244
LMArena Multi-Turn13711257
EQ-Bench Creative Writing—752
WildBench78.9%—

Frequently asked questions

Is GLM-4.5-Air better than GPT-4?

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 29.1 on the Noometry Index.

Which is cheaper, GLM-4.5-Air or GPT-4?

GLM-4.5-Air is cheaper. It lists at $0.20 per million input tokens and $1.10 per million output tokens; GPT-4 lists at $30 and $60.

Is GLM-4.5-Air or GPT-4 better for coding?

GLM-4.5-Air scores higher on coding benchmarks: 33.3 versus 31.6 in the Noometry coding category.

Which has the bigger context window?

GLM-4.5-Air does, with 131K tokens against 8K.

How many benchmarks do GLM-4.5-Air and GPT-4 share?

18 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper