Z.ai (Zhipu), open weights

# GLM-5.1

> GLM-5.1 by Z.ai (Zhipu), released April 2026. Ranked #59 of 354 with a Noometry Index of 47.8. API: $1.40 in / $4.40 out per M tokens. 200K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/glm-5-1
- Last updated: 2026-10-10
- Title: GLM-5.1 Benchmarks, Price & Rank (October 2026) | Noometry

GLM-5.1 by Z.ai (Zhipu) ranks 59th of 354 ranked models on the Noometry Index as of October 2026, with a score of 47.8. Its strongest category is writing & preference, where it ranks 31st. API pricing starts at $1.40 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #59 of 354
- **Index score:** 47.8
- **Evidence:** Confirmed 41 results
- **Provider:** [Z.ai (Zhipu)](https://noometry.com/providers/zai)
- **Released:** April 7, 2026
- **Weights:** Open weights
- **Reasoning:** Yes
- **Context window:** 200K
- **Max output:** 131K
- **Input price:** $1.40 / M
- **Output price:** $4.40 / M
- **Blended price:** $2.15 / M
- **Output speed:** Not measured
- **Value:** #148 of 219
- **Knowledge cutoff:** April 2025
- **Input:** text
- **Hugging Face:** [zai-org/GLM-5.1](https://huggingface.co/zai-org/GLM-5.1)

## Category scores

Each category score combines every public result we have in that category.

GLM-5.1 category scores

1.  Coding 48.7
2.  Agentic & Tool Use 24.9
3.  Reasoning 39.1
4.  Math 49.7
5.  Knowledge 54.9
6.  Multilingual 55.0
7.  Instruction Following 76.3
8.  Long Context 44.9
9.  Writing & Preference 66.9
10.  020406080

GLM-5.1 category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 48.7 | #55 | 5 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 24.9 | #113 | 3 |
| [Reasoning](https://noometry.com/best/reasoning) | 39.1 | #60 | 6 |
| [Math](https://noometry.com/best/math) | 49.7 | #60 | 5 |
| [Knowledge](https://noometry.com/best/knowledge) | 54.9 | #50 | 3 |
| [Multilingual](https://noometry.com/best/multilingual) | 55.0 | #36 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 76.3 | #42 | 1 |
| [Long Context](https://noometry.com/best/long-context) | 44.9 | #53 | 1 |
| [Writing & Preference](https://noometry.com/best/writing) | 66.9 | #31 | 4 |

## Strengths and weaknesses

Categories where GLM-5.1 places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

GLM-5.1: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Writing & Preference](https://noometry.com/best/writing) | 66.9 | +13.2 | #31 of 312, top 10% |
| [Multilingual](https://noometry.com/best/multilingual) | 55.0 | +7.5 | #36 of 297, top 13% |
| [Instruction Following](https://noometry.com/best/instruction-following) | 76.3 | +5.0 | #42 of 305, top 14% |

### Weakest categories

GLM-5.1: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 24.9 | −5.5 | #113 of 154, top 74% |
| [Math](https://noometry.com/best/math) | 49.7 | +13.1 | #60 of 327, top 19% |
| [Long Context](https://noometry.com/best/long-context) | 44.9 | +3.9 | #53 of 296, top 18% |

## Closest competitors

The models ranked just above and below GLM-5.1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GLM-5.1
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [MiMo-V2.6-Flash](https://noometry.com/models/mimo-v2-6-flash) | #55 | 48.5 | $0.18 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-mimo-v2-6-flash) |
| [Grok 4](https://noometry.com/models/grok-4) | #56 | 48.1 | — | 1 | [Compare](https://noometry.com/compare/glm-5-1-vs-grok-4) |
| [Kimi K2.5](https://noometry.com/models/kimi-k2-5) | #57 | 48.1 | $0.90 | 66 | [Compare](https://noometry.com/compare/glm-5-1-vs-kimi-k2-5) |
| [Step 5 Preview](https://noometry.com/models/step-5-preview) | #58 | 47.9 | $1.43 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-step-5-preview) |
| [Kimi K2.6](https://noometry.com/models/kimi-k2-6) | #60 | 47.7 | $1.71 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-kimi-k2-6) |
| [o3](https://noometry.com/models/o3) | #61 | 47.5 | $3.50 | 3 | [Compare](https://noometry.com/compare/glm-5-1-vs-o3) |
| [Qwen3.6 Plus](https://noometry.com/models/qwen3-6-plus) | #62 | 47.5 | $1.13 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-qwen3-6-plus) |
| [Inkling-Small](https://noometry.com/models/inkling-small) | #63 | 46.5 | $0.64 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-inkling-small) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

GLM-5.1 Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified) | 74.2% | #16 of 32, top 50% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-05-15 |
| [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev) | 1508 | #47 of 113, top 42% |  | [LMArena](https://lmarena.ai/leaderboard/webdev) | 2026-10-08 |
| [SciCode](https://noometry.com/benchmarks/scicode) | 43.8% | #62 of 121, top 52% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 57.1% | #36 of 119, top 31% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1485 | #37 of 294, top 13% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [ALE-Bench](https://noometry.com/benchmarks/ale-bench) | 887.1 | #52 of 105, top 50% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Agentic & Tool Use

GLM-5.1 Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [APEX-Agents](https://noometry.com/benchmarks/apex-agents) | 40.9% | #36 of 49, top 74% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [ExploitBench](https://noometry.com/benchmarks/exploitbench) | 18.1% | #7 of 9, top 78% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [GBAEval](https://noometry.com/benchmarks/gbaeval) | 0% | #22 of 23, top 96% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Vending-Bench 2](https://noometry.com/benchmarks/vending-bench-2) | 5,634 | #23 of 60, top 39% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

GLM-5.1 Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [SimpleBench](https://noometry.com/benchmarks/simplebench) | 55.1% | #34 of 77, top 45% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections) | 77.7% | #40 of 91, top 44% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/nyt-connections) |  |
| [CritPt](https://noometry.com/benchmarks/critpt) | 4.6% | #59 of 134, top 45% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 19% | #62 of 129, top 49% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-10 |
| [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization) | 69.8% | #6 of 23, top 27% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1472 | #37 of 297, top 13% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 149.84 | #55 of 213, top 26% |  | [Epoch AI](https://epoch.ai/eci) | 2026-04-07 |

### Math

GLM-5.1 Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [FrontierMath (Tiers 1-3)](https://noometry.com/benchmarks/frontiermath) | 36.8% | #54 of 81, top 67% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-29 |
| [FrontierMath (Tiers 1-3)](https://noometry.com/benchmarks/frontiermath) | 24.9% |  | none | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |
| [MathArena Final-Answer Competitions](https://noometry.com/benchmarks/matharena) | 67.1% | #17 of 29, top 59% |  | [MathArena](https://matharena.ai/) |  |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 93.3% | #41 of 173, top 24% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-10 |
| [ProofBench](https://noometry.com/benchmarks/proofbench) | 22.2% | #45 of 77, top 59% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1473 | #36 of 285, top 13% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [FrontierMath (Feb 2025 set)](https://noometry.com/benchmarks/frontiermath-2025-02) | 33.4% | #15 of 68, top 23% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-05-11 |
| [FrontierMath Tier 4 (v1)](https://noometry.com/benchmarks/frontiermath-tier-4-v1) | 12.5% | #17 of 55, top 31% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-05-12 |

### Knowledge

GLM-5.1 Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 89.9% | #39 of 186, top 21% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-10 |
| [SimpleQA Verified](https://noometry.com/benchmarks/simpleqa-verified) | 34% | #51 of 77, top 67% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-27 |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1476 | #46 of 273, top 17% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Multilingual

GLM-5.1 Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1447 | #36 of 297, top 13% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1515 | #30 of 285, top 11% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1474 | #32 of 223, top 15% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1465 | #31 of 231, top 14% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1434 | #31 of 211, top 15% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1418 | #36 of 213, top 17% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1454 | #39 of 283, top 14% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1469 | #27 of 226, top 12% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

GLM-5.1 Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1451 | #40 of 298, top 14% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

GLM-5.1 Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1466 | #35 of 291, top 13% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

GLM-5.1 Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1461 | #33 of 297, top 12% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1453 | #27 of 295, top 10% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [EQ-Bench Creative Writing](https://noometry.com/benchmarks/eqbench-creative-writing) | 1592 | #43 of 115, top 38% |  | [EQ-Bench](https://eqbench.com/creative_writing.html) |  |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1472 | #27 of 295, top 10% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## API pricing by provider

GLM-5.1 API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [deepinfra](https://deepinfra.com/models) | $1.05 | $3.50 | $0.20 | 2026-10-10 |
| [openrouter](https://openrouter.ai/z-ai/glm-5.1) | $0.97 | $3.04 | $0.18 | 2026-10-10 |
| [together](https://docs.together.ai/docs/serverless-models) | $1.40 | $4.40 | $0.26 | 2026-10-10 |
| [zai](https://docs.z.ai/guides/overview/pricing) | $1.40 | $4.40 | $0.26 | 2026-10-10 |

[All Z.ai (Zhipu) API prices →](https://noometry.com/llm-pricing/zai) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare GLM-5.1

-   [GLM-5.1 vs GLM-5](https://noometry.com/compare/glm-5-vs-glm-5-1)
-   [GLM-5.1 vs Step 5 Preview](https://noometry.com/compare/glm-5-1-vs-step-5-preview)
-   [GLM-5.1 vs Kimi K2.6](https://noometry.com/compare/glm-5-1-vs-kimi-k2-6)
-   [GLM-5.1 vs Kimi K2.5](https://noometry.com/compare/glm-5-1-vs-kimi-k2-5)
-   [GLM-5.1 vs o3](https://noometry.com/compare/glm-5-1-vs-o3)
-   [GLM-5.1 vs Grok 4](https://noometry.com/compare/glm-5-1-vs-grok-4)
-   [GLM-5.1 vs Qwen3.6 Plus](https://noometry.com/compare/glm-5-1-vs-qwen3-6-plus)
-   [GLM-5.1 vs GPT-6 Astra](https://noometry.com/compare/glm-5-1-vs-gpt-6-astra)
-   [GLM-5.1 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-glm-5-1)
-   [GLM-5.1 vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-glm-5-1)
-   [GLM-5.1 vs Kimi K3](https://noometry.com/compare/glm-5-1-vs-kimi-k3)
-   [GLM-5.1 vs Grok 4.6](https://noometry.com/compare/glm-5-1-vs-grok-4-6)
-   [GLM-5.1 vs Qwen3.8 Max](https://noometry.com/compare/glm-5-1-vs-qwen3-8-max)
-   [GLM-5.1 vs Muse Spark 1.3](https://noometry.com/compare/glm-5-1-vs-muse-spark-1-3)

## Other Z.ai (Zhipu) models

-   [GLM-5.3](https://noometry.com/models/glm-5-3)54.8
-   [GLM-5.3-Flash](https://noometry.com/models/glm-5-3-flash)51.8
-   [GLM-5.2](https://noometry.com/models/glm-5-2)51.1
-   [GLM-5](https://noometry.com/models/glm-5)46.1
-   [GLM-5V-Turbo](https://noometry.com/models/glm-5v-turbo)43.8
-   [GLM-4.5](https://noometry.com/models/glm-4-5)42.0
-   [GLM-4.7](https://noometry.com/models/glm-4-7)42.0
-   [GLM-4.6](https://noometry.com/models/glm-4-6)41.4

## Frequently asked questions

### How good is GLM-5.1?

GLM-5.1 by Z.ai (Zhipu) ranks 59th of 354 ranked models on the Noometry Index as of October 2026, with a score of 47.8. Its strongest category is writing & preference, where it ranks 31st. API pricing starts at $1.40 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.

### How much does GLM-5.1 cost?

GLM-5.1 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai (Zhipu)'s own API, with cached input at $0.26.

### What is GLM-5.1's context window?

GLM-5.1 accepts up to 200K tokens of input and can write up to 131K tokens in one response.

### Is GLM-5.1 open source?

Yes. GLM-5.1's weights are downloadable from Hugging Face (zai-org/GLM-5.1); check the license for commercial terms.

### What are GLM-5.1's strengths and weaknesses?

Relative to other ranked models, GLM-5.1 places best in writing & preference, multilingual, instruction following and lowest in agentic & tool use, math, long context.

### What is GLM-5.1 best at?

Its best category is writing & preference, where it ranks 31st on Noometry.

### Cite this page

Noometry. (2026). GLM-5.1 benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/glm-5-1

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/glm-5-1.md).
