Z.ai (Zhipu), open weights
GLM-5.3
GLM-5.3 by Z.ai (Zhipu) ranks 26th of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.8. Its strongest category is writing & preference, where it ranks 6th. API pricing starts at $1.40 per million input tokens and $4.40 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #26 of 354
- Index score
- 54.8
- Evidence
- Confirmed 42 results
- Provider
- Z.ai (Zhipu)
- Released
- August 14, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 1M
- Max output
- 131K
- Input price
- $1.40 / M
- Output price
- $4.40 / M
- Blended price
- $2.15 / M
- Output speed
- Not measured
- Value
- #141 of 219
- Knowledge cutoff
- Unknown
- Input
- text
- Hugging Face
- zai-org/GLM-5.3
Category scores
Each category score combines every public result we have in that category.
- Coding 59.5
- Agentic & Tool Use 36.4
- Reasoning 46.1
- Math 62.3
- Knowledge 58.3
- Multilingual 55.7
- Instruction Following 77.5
- Long Context 45.4
- Writing & Preference 75.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 59.5 | #14 | 8 |
| Agentic & Tool Use | 36.4 | #38 | 1 |
| Reasoning | 46.1 | #46 | 7 |
| Math | 62.3 | #33 | 5 |
| Knowledge | 58.3 | #37 | 3 |
| Multilingual | 55.7 | #28 | 1 |
| Instruction Following | 77.5 | #23 | 1 |
| Long Context | 45.4 | #41 | 1 |
| Writing & Preference | 75.7 | #6 | 4 |
Strengths and weaknesses
Categories where GLM-5.3 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 75.7 | +21.9 | #6 of 312, top 2% |
| Coding | 59.5 | +20.8 | #14 of 340, top 5% |
| Instruction Following | 77.5 | +6.2 | #23 of 305, top 8% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 36.4 | +6.0 | #38 of 154, top 25% |
| Long Context | 45.4 | +4.5 | #41 of 296, top 14% |
| Reasoning | 46.1 | +22.5 | #46 of 350, top 14% |
Closest competitors
The models ranked just above and below GLM-5.3. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen3.8 Max | #22 | 56.8 | $3 | — | Compare |
| Gemini 3.1 Pro Preview | #23 | 56.7 | $4.50 | — | Compare |
| Gemini 4 Argon | #24 | 56.5 | — | — | Compare |
| Grok 4.5 | #25 | 55.0 | $3 | 4 | Compare |
| Muse Spark 1.3 | #27 | 54.8 | $2 | — | Compare |
| Gemini 3 Pro | #28 | 54.8 | — | 1 | Compare |
| Claude Sonnet 5 | #29 | 54.6 | $4 | — | Compare |
| GPT-5.6 Luna | #30 | 54.6 | $0.45 | 12 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| DeepSWE | 69% | #8 of 29, top 28% | max | Epoch AI | |
| FrontierCode | 40.1% | #21 of 37, top 57% | max | Epoch AI | |
| CursorBench | 38% | high | Epoch AI | ||
| CursorBench | 33.3% | low | Epoch AI | ||
| CursorBench | 42.6% | #6 of 14, top 43% | max | Epoch AI | |
| LMArena WebDev | 1622 | #17 of 113, top 16% | LMArena | 2026-10-08 | |
| FrontierSWE | 30.2% | #9 of 18, top 50% | max | Epoch AI | |
| SciCode | 42% | low | Epoch AI | ||
| SciCode | 59% | #10 of 121, top 9% | max | Epoch AI | |
| WeirdML | 75.4% | #15 of 119, top 13% | max | Epoch AI | |
| LMArena Coding | 1496 | #25 of 294, top 9% | LMArena | 2026-10-08 | |
| ALE-Bench | 1,317 | #23 of 105, top 22% | high | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 56.6% | #16 of 49, top 33% | Epoch AI | ||
| Vending-Bench 2 | 8,164 | #11 of 60, top 19% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| NYT Connections (extended) | 74.2% | #46 of 91, top 51% | high reasoning | Lech Mazur benchmarks | |
| CritPt | 14.6% | low | Epoch AI | ||
| CritPt | 19.1% | #27 of 134, top 21% | max | Epoch AI | |
| Chess Puzzles | 21% | #53 of 129, top 42% | max | Epoch AI | 2026-08-24 |
| LMArena Hard Prompts | 1489 | #16 of 297, top 6% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 33% | #23 of 74, top 32% | max | Epoch AI | 2026-08-30 |
| DTBench | 87.7% | #49 of 151, top 33% | Epoch AI | ||
| LMCA | 55.5% | #10 of 125, top 8% | Epoch AI | ||
| Bench to the Future 3 | 0.15 | Best of 10 | Epoch AI | ||
| Epoch Capabilities Index | 155.61 | #27 of 213, top 13% | Epoch AI | 2026-08-14 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 68.8% | #25 of 81, top 31% | max | Epoch AI | 2026-08-25 |
| FrontierMath Tier 4 | 29.3% | #31 of 63, top 50% | max | Epoch AI | 2026-08-25 |
| OTIS Mock AIME 2024-2025 | 91.1% | #49 of 173, top 29% | max | Epoch AI | 2026-08-24 |
| ProofBench | 49% | #31 of 77, top 41% | max | Epoch AI | |
| LMArena Math | 1489 | #19 of 285, top 7% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 90.9% | #29 of 186, top 16% | max | Epoch AI | 2026-08-24 |
| SimpleQA Verified | 41% | #41 of 77, top 54% | max | Epoch AI | 2026-08-28 |
| LMArena Expert | 1516 | #14 of 273, top 6% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1457 | #28 of 297, top 10% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1528 | #19 of 285, top 7% | LMArena | 2026-10-08 | |
| LMArena French | 1499 | #12 of 223, top 6% | LMArena | 2026-10-08 | |
| LMArena German | 1499 | #6 of 231, top 3% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1453 | #22 of 211, top 11% | LMArena | 2026-10-08 | |
| LMArena Korean | 1472 | #6 of 213, top 3% | LMArena | 2026-10-08 | |
| LMArena Russian | 1463 | #31 of 283, top 11% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1460 | #35 of 226, top 16% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1477 | #20 of 298, top 7% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1482 | #22 of 291, top 8% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1471 | #24 of 297, top 9% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1457 | #20 of 295, top 7% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 2075 | #6 of 115, top 6% | EQ-Bench | ||
| LMArena Multi-Turn | 1472 | #28 of 295, top 10% | LMArena | 2026-10-08 |
API pricing by provider
Compare GLM-5.3
- GLM-5.3 vs GLM-5.2
- GLM-5.3 vs Grok 4.5
- GLM-5.3 vs Muse Spark 1.3
- GLM-5.3 vs Gemini 4 Argon
- GLM-5.3 vs Gemini 3 Pro
- GLM-5.3 vs Gemini 3.1 Pro Preview
- GLM-5.3 vs Claude Sonnet 5
- GLM-5.3 vs GPT-6 Astra
- GLM-5.3 vs Claude Fable 5.1
- GLM-5.3 vs Gemini 3.8 Flash
- GLM-5.3 vs Kimi K3
- GLM-5.3 vs Grok 4.6
- GLM-5.3 vs Qwen3.8 Max
- GLM-5.3 vs DeepSeek V4 Pro
Other Z.ai (Zhipu) models
- GLM-5.3-Flash51.8
- GLM-5.251.1
- GLM-5.147.8
- GLM-546.1
- GLM-5V-Turbo43.8
- GLM-4.542.0
- GLM-4.742.0
- GLM-4.641.4
Frequently asked questions
How good is GLM-5.3?
GLM-5.3 by Z.ai (Zhipu) ranks 26th of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.8. Its strongest category is writing & preference, where it ranks 6th. API pricing starts at $1.40 per million input tokens and $4.40 per million output tokens, with a 1M-token context window.
How much does GLM-5.3 cost?
GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai (Zhipu)'s own API, with cached input at $0.26.
What is GLM-5.3's context window?
GLM-5.3 accepts up to 1M tokens of input and can write up to 131K tokens in one response.
Is GLM-5.3 open source?
Yes. GLM-5.3's weights are downloadable from Hugging Face (zai-org/GLM-5.3); check the license for commercial terms.
What are GLM-5.3's strengths and weaknesses?
Relative to other ranked models, GLM-5.3 places best in writing & preference, coding, instruction following and lowest in agentic & tool use, long context, reasoning.
What is GLM-5.3 best at?
Its best category is writing & preference, where it ranks 6th on Noometry.