Model comparison

Claude 2 vs GLM-4.5-Air

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 25.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Summary

  • The widest gap is in math, where GLM-4.5-Air leads 36.2 to 9.3.
  • GLM-4.5-Air has downloadable open weights; the other is API-only.

Side by side

Claude 2 and GLM-4.5-Air specifications
Claude 2GLM-4.5-Air
ProviderAnthropicZ.ai (Zhipu)
Noometry Index25.038.9
Released2023-07-112025-07-20
WeightsProprietaryOpen
Context window—131K
Max output—98K
Input $ / M tokens—$0.20
Output $ / M tokens—$1.10
Results tracked827

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, GLM-4.5-Air: 33.3 (#259)

Coding benchmarks
BenchmarkClaude 2GLM-4.5-Air
GSO—2.9%
LMArena Coding—1397
HumanEval+61.6%—

Reasoning GLM-4.5-Air leads

Claude 2: 21.7 (#216), GLM-4.5-Air: 24.1 (#166)

Reasoning benchmarks
BenchmarkClaude 2GLM-4.5-Air
Kagi LLM Benchmark—43%
LMArena Hard Prompts—1379
DTBench51.9%—
Epoch Capabilities Index120.13—
ForecastBench—59.2

Math GLM-4.5-Air leads

Claude 2: 9.3 (#320), GLM-4.5-Air: 36.2 (#170)

Math benchmarks
BenchmarkClaude 2GLM-4.5-Air
OTIS Mock AIME 2024-20252.5%—
Omni-MATH—39.1%
LMArena Math—1396
MATH Level 511.7%—

Knowledge GLM-4.5-Air leads

Claude 2: 16.9 (#287), GLM-4.5-Air: 35.0 (#191)

Knowledge benchmarks
BenchmarkClaude 2GLM-4.5-Air
GPQA Diamond34.7%—
Humanity's Last Exam—8.1%
MMLU-Pro—76.2%
Vectara Hallucination Rate—9.3%
GPQA (HELM)—59.4%
LMArena Expert—1370
MMLU78.5%—
TriviaQA87.5%—

Multilingual Not comparable

Claude 2: —, GLM-4.5-Air: 49.1 (#135)

Multilingual benchmarks
BenchmarkClaude 2GLM-4.5-Air
LMArena Non-English—1366
LMArena Chinese—1426
LMArena French—1399
LMArena German—1377
LMArena Japanese—1348
LMArena Korean—1308
LMArena Russian—1373
LMArena Spanish—1386

Instruction Following Not comparable

Claude 2: —, GLM-4.5-Air: 69.6 (#171)

Instruction Following benchmarks
BenchmarkClaude 2GLM-4.5-Air
IFEval—81.2%
LMArena Instruction Following—1354

Long Context Not comparable

Claude 2: —, GLM-4.5-Air: 41.6 (#135)

Long Context benchmarks
BenchmarkClaude 2GLM-4.5-Air
LMArena Longer Query—1366

Writing & Preference Not comparable

Claude 2: —, GLM-4.5-Air: 55.9 (#139)

Writing & Preference benchmarks
BenchmarkClaude 2GLM-4.5-Air
LMArena Text—1384
LMArena Creative Writing—1343
WildBench—78.9%
LMArena Multi-Turn—1371

Frequently asked questions

Is Claude 2 better than GLM-4.5-Air?

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and GLM-4.5-Air share?

0 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and GLM-4.5-Air has 27.

Related comparisons

Go deeper