Model comparison

Claude Sonnet 4.5 vs GLM-5

GLM-5 is the stronger model overall, scoring 46.1 to 44.1 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Sonnet 4.5 Anthropic

44.1

Rank #81 Confirmed

GLM-5 Z.ai (Zhipu)

46.1

Rank #66 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Sonnet 4.5 scores higher in 3 categories and GLM-5 in 6 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5 leads 46.4 to 32.3.
  • The biggest single-benchmark swing is NYT Connections (extended): 37.3% for Claude Sonnet 4.5 and 74.8% for GLM-5.
  • GLM-5 is cheaper at $1 / $3.20 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.5.
  • GLM-5 accepts more context: 205K tokens versus 200K.
  • GLM-5 has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4.5 and GLM-5 specifications
Claude Sonnet 4.5GLM-5
ProviderAnthropicZ.ai (Zhipu)
Noometry Index44.146.1
Released2025-09-292026-02-11
WeightsProprietaryOpen
Context window200K205K
Max output64K131K
Input $ / M tokens$3$1
Output $ / M tokens$15$3.20
Results tracked7345

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5 leads

Claude Sonnet 4.5: 47.3 (#61), GLM-5: 49.0 (#52)

Coding benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
SWE-bench Verified71.3%72.1%
SWE-bench Verified (bash only)71.4%72.8%
LMArena WebDev13931434
SWE-bench Multilingual67%69.7%
WeirdML47.7%48.2%
LMArena Coding14891461
ALE-Bench796.15765.62
SciCode44.7%—
GSO14.7%—
AlgoTune1.52—

Agentic & Tool Use Claude Sonnet 4.5 leads

Claude Sonnet 4.5: 38.3 (#32), GLM-5: 31.1 (#71)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
Terminal-Bench46.5%52.4%
τ²-bench Airline72%82.5%
τ²-bench Banking25.3%9.8%
τ²-bench Retail72.4%73.7%
τ²-bench Telecom84.9%86.8%
Vending-Bench 23,8394,432
Berkeley Function Calling Leaderboard73.2%—
GDPval42.5%—
Remote Labor Index2.1%—
Cybench60%—
DeepResearch Bench52.6%—
OSWorld62.9%—
LMArena Search1159—
METR Time Horizons67.4%—

Reasoning Too close to call

Claude Sonnet 4.5: 26.9 (#125), GLM-5: 27.6 (#116)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
ARC-AGI-213.6%4.9%
SimpleBench54.3%53.2%
Kagi LLM Benchmark57.9%75%
NYT Connections (extended)37.3%74.8%
ARC-AGI-163.7%44.7%
Chess Puzzles12%10%
LMArena Hard Prompts14621452
Epoch Capabilities Index146.84145.83
ForecastBench61.961
CritPt1.1%—
EnigmaEval6%—
EBR-Bench2.4%—
Mystery Game Puzzles17%—
DTBench83.2%—
LMCA38.8%—

Math GLM-5 leads

Claude Sonnet 4.5: 32.3 (#216), GLM-5: 46.4 (#71)

Knowledge GLM-5 leads

Claude Sonnet 4.5: 48.4 (#76), GLM-5: 52.3 (#64)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
GPQA Diamond82.3%87.8%
Vectara Hallucination Rate12%10.1%
LMArena Expert14821454
Humanity's Last Exam13.7%—
SimpleQA Verified30.7%—
MMLU-Pro86.9%—
GPQA (HELM)68.6%—

Multimodal Not comparable

Claude Sonnet 4.5: 34.8 (#89), GLM-5: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
VPCT39.8%—
LMArena Document1450—

Multilingual Too close to call

Claude Sonnet 4.5: 53.4 (#69), GLM-5: 53.7 (#58)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
LMArena Non-English14251430
LMArena Chinese14591511
LMArena French14581455
LMArena German14271445
LMArena Japanese13901416
LMArena Korean14031423
LMArena Russian14371436
LMArena Spanish14571454

Instruction Following Too close to call

Claude Sonnet 4.5: 75.0 (#78), GLM-5: 75.2 (#67)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
LMArena Instruction Following14591428
IFEval85%—

Long Context Too close to call

Claude Sonnet 4.5: 45.2 (#46), GLM-5: 44.7 (#60)

Long Context benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
LMArena Longer Query14761446
CL-bench—18.7%

Writing & Preference Too close to call

Claude Sonnet 4.5: 66.5 (#34), GLM-5: 66.0 (#38)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.5GLM-5
LMArena Text14391446
LMArena Creative Writing14421439
EQ-Bench Creative Writing16781601
LMArena Multi-Turn14651456
WildBench85.4%—

Frequently asked questions

Is Claude Sonnet 4.5 better than GLM-5?

GLM-5 is the stronger model overall, scoring 46.1 to 44.1 on the Noometry Index.

Which is cheaper, Claude Sonnet 4.5 or GLM-5?

GLM-5 is cheaper. It lists at $1 per million input tokens and $3.20 per million output tokens; Claude Sonnet 4.5 lists at $3 and $15.

Is Claude Sonnet 4.5 or GLM-5 better for coding?

GLM-5 scores higher on coding benchmarks: 49.0 versus 47.3 in the Noometry coding category.

Which has the bigger context window?

GLM-5 does, with 205K tokens against 200K.

How many benchmarks do Claude Sonnet 4.5 and GLM-5 share?

43 benchmarks have published results for both models. Claude Sonnet 4.5 has 73 scored results on Noometry and GLM-5 has 45.

Related comparisons

Go deeper