Z.ai (Zhipu), open weights

GLM-5

GLM-5 by Z.ai (Zhipu) ranks 66th of 354 ranked models on the Noometry Index as of October 2026, with a score of 46.1. Its strongest category is writing & preference, where it ranks 38th. API pricing starts at $1 per million input tokens and $3.20 per million output tokens, with a 205K-token context window.

Last verified

Specifications

Noometry rank
#66 of 354
Index score
46.1
Evidence
Confirmed 45 results
Released
February 11, 2026
Weights
Open weights
Reasoning
Yes
Context window
205K
Max output
131K
Input price
$1 / M
Output price
$3.20 / M
Blended price
$1.55 / M
Output speed
23 tokens/s Kagi
Value
#134 of 219
Knowledge cutoff
December 2025
Input
text
Hugging Face
zai-org/GLM-5

Category scores

Each category score combines every public result we have in that category.

GLM-5 category scores
  1. Coding 49.0
  2. Agentic & Tool Use 31.1
  3. Reasoning 27.6
  4. Math 46.4
  5. Knowledge 52.3
  6. Multilingual 53.7
  7. Instruction Following 75.2
  8. Long Context 44.7
  9. Writing & Preference 66.0
GLM-5 category ranks
CategoryScoreRankResults
Coding49.0#526
Agentic & Tool Use31.1#715
Reasoning27.6#1167
Math46.4#713
Knowledge52.3#643
Multilingual53.7#581
Instruction Following75.2#671
Long Context44.7#602
Writing & Preference66.0#384

Strengths and weaknesses

Categories where GLM-5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GLM-5: strongest categories
CategoryScorevs medianRank
Writing & Preference66.0+12.2#38 of 312, top 13%
Coding49.0+10.3#52 of 340, top 16%
Multilingual53.7+6.3#58 of 297, top 20%

Weakest categories

GLM-5: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use31.1+0.7#71 of 154, top 47%
Reasoning27.6+4.0#116 of 350, top 34%
Instruction Following75.2+3.9#67 of 305, top 22%

Closest competitors

The models ranked just above and below GLM-5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GLM-5
ModelRankScoreBlended $/MSpeed
Qwen3.6 Plus#6247.5$1.13—Compare
Inkling-Small#6346.5$0.64—Compare
GPT-5 Pro#6446.4$41.255Compare
Grok 4.20 Multi-Agent#6546.2$1.56—Compare
Qwen3.5 397B-A17B#6746.0$1.359Compare
Qwen3.8 27B#6846.0$1.11—Compare
GPT-5.3 Codex#6945.8$4.81—Compare
Kimi K2 Thinking Turbo#7045.8——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GLM-5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified72.1%#22 of 32, top 69%Epoch AI2026-02-15
SWE-bench Verified (bash only)72.8%#6 of 39, top 16%highSWE-bench2026-02-17
LMArena WebDev1434#63 of 113, top 56%LMArena2026-10-08
SWE-bench Multilingual69.7%#4 of 13, top 31%SWE-bench2026-02-13
WeirdML48.2%#50 of 119, top 43%Epoch AI
LMArena Coding1461#69 of 294, top 24%LMArena2026-10-08
ALE-Bench765.62#63 of 105, top 60%Epoch AI

Agentic & Tool Use

GLM-5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench52.4%#16 of 41, top 40%Epoch AI
τ²-bench Airline82.5%#4 of 7, top 58%enabledτ²-bench2026-03-02
τ²-bench Banking9.8%#25 of 26, top 97%enabledτ²-bench2026-03-02
τ²-bench Retail73.7%#6 of 7, top 86%enabledτ²-bench2026-03-02
τ²-bench Telecom86.8%#6 of 7, top 86%enabledτ²-bench2026-03-02
Vending-Bench 24,432#33 of 60, top 56%Epoch AI

Reasoning

GLM-5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-24.9%#57 of 83, top 69%Epoch AI
SimpleBench53.2%#36 of 77, top 47%Epoch AI
Kagi LLM Benchmark75%#12 of 99, top 13%Kagi LLM Benchmark
Kagi LLM Benchmark51.7%Kagi LLM Benchmark
NYT Connections (extended)74.8%#43 of 91, top 48%Lech Mazur benchmarks
ARC-AGI-144.7%#59 of 83, top 72%Epoch AI
Chess Puzzles10%#82 of 129, top 64%Epoch AI2026-02-12
LMArena Hard Prompts1452#58 of 297, top 20%LMArena2026-10-08
Epoch Capabilities Index145.83#78 of 213, top 37%Epoch AI2026-02-11
ForecastBench61#18 of 72, top 25%Epoch AI

Math

GLM-5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
MathArena Final-Answer Competitions65.7%#20 of 29, top 69%MathArena
OTIS Mock AIME 2024-202580%#81 of 173, top 47%Epoch AI2026-02-12
LMArena Math1440#73 of 285, top 26%LMArena2026-10-08
FrontierMath (Feb 2025 set)16.4%#33 of 68, top 49%Epoch AI2026-02-19
FrontierMath Tier 4 (v1)2.1%#38 of 55, top 70%Epoch AI2026-02-19

Knowledge

GLM-5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond87.8%#51 of 186, top 28%Epoch AI2026-02-12
Vectara Hallucination Rate (lower is better)10.1%#55 of 96, top 58%Vectara Hallucination Leaderboard
LMArena Expert1454#67 of 273, top 25%LMArena2026-10-08

Multilingual

GLM-5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1430#58 of 297, top 20%LMArena2026-10-08
LMArena Chinese1511#35 of 285, top 13%LMArena2026-10-08
LMArena French1455#60 of 223, top 27%LMArena2026-10-08
LMArena German1445#52 of 231, top 23%LMArena2026-10-08
LMArena Japanese1416#43 of 211, top 21%LMArena2026-10-08
LMArena Korean1423#35 of 213, top 17%LMArena2026-10-08
LMArena Russian1436#56 of 283, top 20%LMArena2026-10-08
LMArena Spanish1454#48 of 226, top 22%LMArena2026-10-08

Instruction Following

GLM-5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1428#63 of 298, top 22%LMArena2026-10-08

Long Context

GLM-5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench18.7%#10 of 19, top 53%Epoch AI
LMArena Longer Query1446#57 of 291, top 20%LMArena2026-10-08

Writing & Preference

GLM-5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1446#51 of 297, top 18%LMArena2026-10-08
LMArena Creative Writing1439#39 of 295, top 14%LMArena2026-10-08
EQ-Bench Creative Writing1601#42 of 115, top 37%EQ-Bench
LMArena Multi-Turn1456#42 of 295, top 15%LMArena2026-10-08

API pricing by provider

GLM-5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$1$3.20—2026-10-10
deepinfra$0.60$2.08$0.122026-10-10
openrouter$0.60$1.92$0.122026-10-10
together$1$3.20—2026-10-10
vertex$1$3.20$0.102026-10-10
zai$1$3.20$0.202026-10-10

Compare GLM-5

Other Z.ai (Zhipu) models

Frequently asked questions

How good is GLM-5?

GLM-5 by Z.ai (Zhipu) ranks 66th of 354 ranked models on the Noometry Index as of October 2026, with a score of 46.1. Its strongest category is writing & preference, where it ranks 38th. API pricing starts at $1 per million input tokens and $3.20 per million output tokens, with a 205K-token context window.

How much does GLM-5 cost?

GLM-5 costs $1 per million input tokens and $3.20 per million output tokens on Z.ai (Zhipu)'s own API, with cached input at $0.20.

What is GLM-5's context window?

GLM-5 accepts up to 205K tokens of input and can write up to 131K tokens in one response.

Is GLM-5 open source?

Yes. GLM-5's weights are downloadable from Hugging Face (zai-org/GLM-5); check the license for commercial terms.

How fast is GLM-5?

GLM-5 generated about 23 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GLM-5's strengths and weaknesses?

Relative to other ranked models, GLM-5 places best in writing & preference, coding, multilingual and lowest in agentic & tool use, reasoning, instruction following.

What is GLM-5 best at?

Its best category is writing & preference, where it ranks 38th on Noometry.