Model comparison

Claude Haiku 4.5 vs GLM-4.5-Air

Claude Haiku 4.5 and GLM-4.5-Air score almost the same on the Noometry Index (39.5 vs 38.9), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Claude Haiku 4.5 scores higher in 7 categories and GLM-4.5-Air in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Claude Haiku 4.5 leads 44.0 to 33.3.
  • The biggest single-benchmark swing is Omni-MATH: 56.1% for Claude Haiku 4.5 and 39.1% for GLM-4.5-Air.
  • GLM-4.5-Air is cheaper at $0.20 / $1.10 per million input/output tokens, against $1 / $5 for Claude Haiku 4.5.
  • Claude Haiku 4.5 accepts more context: 200K tokens versus 131K.
  • GLM-4.5-Air has downloadable open weights; the other is API-only.

Side by side

Claude Haiku 4.5 and GLM-4.5-Air specifications
Claude Haiku 4.5GLM-4.5-Air
ProviderAnthropicZ.ai (Zhipu)
Noometry Index39.538.9
Released2025-10-152025-07-20
WeightsProprietaryOpen
Context window200K131K
Max output64K98K
Input $ / M tokens$1$0.20
Output $ / M tokens$5$1.10
Results tracked5327

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.0 (#78), GLM-4.5-Air: 33.3 (#259)

Coding benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
LMArena Coding14531397
SWE-bench Verified (bash only)66.6%—
LMArena WebDev1330—
SWE-bench Multilingual64.7%—
SciCode43.3%—
GSO—2.9%
WeirdML45.4%—
ALE-Bench653.48—

Agentic & Tool Use Not comparable

Claude Haiku 4.5: 33.6 (#52), GLM-4.5-Air: —

Agentic & Tool Use benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
Terminal-Bench35.5%—
Berkeley Function Calling Leaderboard68.7%—
DeepResearch Bench45.5%—
BALROG31.2%—
ExploitBench13.7%—
Vending-Bench 2458.89—

Reasoning GLM-4.5-Air leads

Claude Haiku 4.5: 15.1 (#320), GLM-4.5-Air: 24.1 (#166)

Reasoning benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
LMArena Hard Prompts14201379
ForecastBench61.459.2
ARC-AGI-24%—
Kagi LLM Benchmark—43%
NYT Connections (extended)14.3%—
ARC-AGI-147.7%—
CritPt0%—
Chess Puzzles8%—
DTBench73.6%—
LMCA30.9%—
Epoch Capabilities Index142.41—

Math Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.9 (#78), GLM-4.5-Air: 36.2 (#170)

Math benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
Omni-MATH56.1%39.1%
LMArena Math13961396
OTIS Mock AIME 2024-202566.7%—
MATH Level 596.4%—
FrontierMath (Feb 2025 set)5.9%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Claude Haiku 4.5 leads

Claude Haiku 4.5: 37.7 (#153), GLM-4.5-Air: 35.0 (#191)

Knowledge benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
MMLU-Pro77.7%76.2%
Vectara Hallucination Rate9.8%9.3%
GPQA (HELM)60.5%59.4%
LMArena Expert14421370
GPQA Diamond71.2%—
Humanity's Last Exam—8.1%
SimpleQA Verified13.2%—

Multimodal Not comparable

Claude Haiku 4.5: 26.8 (#118), GLM-4.5-Air: —

Multimodal benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
Blueprint-Bench 20%—
LMArena Document1420—

Multilingual Too close to call

Claude Haiku 4.5: 49.9 (#129), GLM-4.5-Air: 49.1 (#135)

Multilingual benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
LMArena Non-English13771366
LMArena Chinese14171426
LMArena French14081399
LMArena German13751377
LMArena Japanese13391348
LMArena Korean13471308
LMArena Russian13811373
LMArena Spanish14201386

Instruction Following Claude Haiku 4.5 leads

Claude Haiku 4.5: 71.4 (#149), GLM-4.5-Air: 69.6 (#171)

Instruction Following benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
IFEval80.1%81.2%
LMArena Instruction Following14141354

Long Context Claude Haiku 4.5 leads

Claude Haiku 4.5: 43.6 (#92), GLM-4.5-Air: 41.6 (#135)

Long Context benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
LMArena Longer Query14271366

Writing & Preference Claude Haiku 4.5 leads

Claude Haiku 4.5: 57.9 (#123), GLM-4.5-Air: 55.9 (#139)

Writing & Preference benchmarks
BenchmarkClaude Haiku 4.5GLM-4.5-Air
LMArena Text13961384
LMArena Creative Writing13721343
WildBench83.9%78.9%
LMArena Multi-Turn14091371
EQ-Bench 41064—

Frequently asked questions

Is Claude Haiku 4.5 better than GLM-4.5-Air?

Claude Haiku 4.5 and GLM-4.5-Air score almost the same on the Noometry Index (39.5 vs 38.9), so choose on price, context window or the category you care about most.

Which is cheaper, Claude Haiku 4.5 or GLM-4.5-Air?

GLM-4.5-Air is cheaper. It lists at $0.20 per million input tokens and $1.10 per million output tokens; Claude Haiku 4.5 lists at $1 and $5.

Is Claude Haiku 4.5 or GLM-4.5-Air better for coding?

Claude Haiku 4.5 scores higher on coding benchmarks: 44.0 versus 33.3 in the Noometry coding category.

Which has the bigger context window?

Claude Haiku 4.5 does, with 200K tokens against 131K.

How many benchmarks do Claude Haiku 4.5 and GLM-4.5-Air share?

24 benchmarks have published results for both models. Claude Haiku 4.5 has 53 scored results on Noometry and GLM-4.5-Air has 27.

Related comparisons

Go deeper