Model comparison
GLM-5.2 vs Qwen3.7 Max
GLM-5.2 and Qwen3.7 Max score almost the same on the Noometry Index (51.1 vs 51.5), so choose on price, context window or the category you care about most.
Last verified . 33 shared benchmarks.
Summary
- They share 33 benchmarks with published results for both. GLM-5.2 scores higher in 4 categories and Qwen3.7 Max in 5 categories; 6 gaps are clear of the uncertainty.
- The widest gap is in agentic & tool use, where GLM-5.2 leads 32.4 to 22.1.
- The biggest single-benchmark swing is SimpleQA Verified: 34.2% for GLM-5.2 and 55.8% for Qwen3.7 Max.
- GLM-5.2 is cheaper at $1.40 / $4.40 per million input/output tokens, against $2.50 / $7.50 for Qwen3.7 Max.
- GLM-5.2 has downloadable open weights; the other is API-only.
Side by side
| GLM-5.2 | Qwen3.7 Max | |
|---|---|---|
| Provider | Z.ai (Zhipu) | Alibaba (Qwen) |
| Noometry Index | 51.1 | 51.5 |
| Released | 2026-06-13 | 2026-05-19 |
| Weights | Open | Proprietary |
| Context window | 1M | 1M |
| Max output | 131K | 131K |
| Input $ / M tokens | $1.40 | $2.50 |
| Output $ / M tokens | $4.40 | $7.50 |
| Results tracked | 51 | 33 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Too close to call
GLM-5.2: 51.3 (#41), Qwen3.7 Max: 50.4 (#45)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| SWE-bench Verified | 78.7% | 77.3% |
| LMArena WebDev | 1603 | 1515 |
| SciCode | 50.5% | 48.8% |
| LMArena Coding | 1485 | 1498 |
| ALE-Bench | 1,047 | 1,189 |
| DeepSWE | 43.8% | — |
| FrontierCode | 24.5% | — |
| WeirdML | 70.1% | — |
Agentic & Tool Use GLM-5.2 leads
GLM-5.2: 32.4 (#63), Qwen3.7 Max: 22.1 (#135)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| GBAEval | 0% | 0.4% |
| APEX-Agents | 45.2% | — |
| τ²-bench Banking | 37.1% | — |
| PostTrainBench | 31.7% | — |
| Vending-Bench 2 | 8,314 | — |
Reasoning Qwen3.7 Max leads
GLM-5.2: 42.3 (#52), Qwen3.7 Max: 49.2 (#38)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| SimpleBench | 58.8% | 70.4% |
| NYT Connections (extended) | 74.3% | 85.1% |
| CritPt | 20.9% | 13.4% |
| Chess Puzzles | 21% | 19% |
| EBR-Bench | 9.5% | 9.5% |
| LMArena Hard Prompts | 1480 | 1483 |
| Mystery Game Puzzles | 19% | 32% |
| DTBench | 93.6% | 92.3% |
| LMCA | 45.8% | 44% |
| Epoch Capabilities Index | 151.78 | 153.68 |
| ARC-AGI-2 | 22.8% | — |
| Kagi LLM Benchmark | 62.6% | — |
| ARC-AGI-1 | 77% | — |
| Surface Evolver Bench | 55.6% | — |
Math Qwen3.7 Max leads
GLM-5.2: 55.7 (#43), Qwen3.7 Max: 62.4 (#32)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| FrontierMath (Tiers 1-3) | 59.2% | 64.6% |
| FrontierMath Tier 4 | 29.3% | 34.1% |
| OTIS Mock AIME 2024-2025 | 86.4% | 95.6% |
| ProofBench | 35% | 26% |
| LMArena Math | 1482 | 1490 |
| MathArena Final-Answer Competitions | 67.6% | — |
Knowledge Qwen3.7 Max leads
GLM-5.2: 57.1 (#40), Qwen3.7 Max: 61.6 (#28)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| GPQA Diamond | 91.9% | 90.9% |
| SimpleQA Verified | 34.2% | 55.8% |
| LMArena Expert | 1486 | 1488 |
Multilingual Qwen3.7 Max leads
GLM-5.2: 55.8 (#26), Qwen3.7 Max: 56.9 (#15)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| LMArena Non-English | 1459 | 1474 |
| LMArena Chinese | 1519 | 1530 |
| LMArena Russian | 1466 | 1484 |
| LMArena French | 1479 | — |
| LMArena German | 1468 | — |
| LMArena Japanese | 1451 | — |
| LMArena Korean | 1445 | — |
| LMArena Spanish | 1477 | — |
Instruction Following Too close to call
GLM-5.2: 76.9 (#34), Qwen3.7 Max: 76.7 (#38)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| LMArena Instruction Following | 1465 | 1460 |
Long Context Too close to call
GLM-5.2: 45.3 (#43), Qwen3.7 Max: 45.4 (#40)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| LMArena Longer Query | 1479 | 1482 |
Writing & Preference GLM-5.2 leads
GLM-5.2: 70.4 (#21), Qwen3.7 Max: 65.0 (#54)
| Benchmark | GLM-5.2 | Qwen3.7 Max |
|---|---|---|
| LMArena Text | 1470 | 1476 |
| LMArena Creative Writing | 1462 | 1449 |
| EQ-Bench 4 | 1222 | 1110 |
| LMArena Multi-Turn | 1469 | 1481 |
| EQ-Bench Creative Writing | 1757 | — |
Frequently asked questions
Is GLM-5.2 better than Qwen3.7 Max?
GLM-5.2 and Qwen3.7 Max score almost the same on the Noometry Index (51.1 vs 51.5), so choose on price, context window or the category you care about most.
Which is cheaper, GLM-5.2 or Qwen3.7 Max?
GLM-5.2 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; Qwen3.7 Max lists at $2.50 and $7.50.
Is GLM-5.2 or Qwen3.7 Max better for coding?
They score almost the same on coding (51.3 vs 50.4); test both on your own repository before choosing.
Which has the bigger context window?
Both accept 1M tokens.
How many benchmarks do GLM-5.2 and Qwen3.7 Max share?
33 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Qwen3.7 Max has 33.