Model comparison

GLM-5 vs GPT-5.5 Instant

GLM-5 is the stronger model overall, scoring 46.1 to 42.7 on the Noometry Index.

Last verified . 21 shared benchmarks.

GLM-5 Z.ai (Zhipu)

46.1

Rank #66 Confirmed

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GLM-5 scores higher in 8 categories and GPT-5.5 Instant in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5 leads 46.4 to 26.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 80% for GLM-5 and 68.1% for GPT-5.5 Instant.
  • GLM-5 has downloadable open weights; the other is API-only.

Side by side

GLM-5 and GPT-5.5 Instant specifications
GLM-5GPT-5.5 Instant
ProviderZ.ai (Zhipu)OpenAI
Noometry Index46.142.7
Released2026-02-112026-05-05
WeightsOpenProprietary
Context window205K—
Max output131K—
Input $ / M tokens$1—
Output $ / M tokens$3.20—
Results tracked4527

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5 leads

GLM-5: 49.0 (#52), GPT-5.5 Instant: 44.3 (#74)

Coding benchmarks
BenchmarkGLM-5GPT-5.5 Instant
LMArena Coding14611433
SWE-bench Verified72.1%—
SWE-bench Verified (bash only)72.8%—
LMArena WebDev1434—
SWE-bench Multilingual69.7%—
SciCode—48.6%
WeirdML48.2%—
ALE-Bench765.62—

Agentic & Tool Use Not comparable

GLM-5: 31.1 (#71), GPT-5.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGLM-5GPT-5.5 Instant
Terminal-Bench52.4%—
τ²-bench Airline82.5%—
τ²-bench Banking9.8%—
τ²-bench Retail73.7%—
τ²-bench Telecom86.8%—
Vending-Bench 24,432—

Reasoning GLM-5 leads

GLM-5: 27.6 (#116), GPT-5.5 Instant: 24.9 (#155)

Reasoning benchmarks
BenchmarkGLM-5GPT-5.5 Instant
Chess Puzzles10%12%
LMArena Hard Prompts14521426
Epoch Capabilities Index145.83142.52
ARC-AGI-24.9%—
SimpleBench53.2%—
Kagi LLM Benchmark75%—
NYT Connections (extended)74.8%—
ARC-AGI-144.7%—
CritPt—0%
ForecastBench61—

Math GLM-5 leads

GLM-5: 46.4 (#71), GPT-5.5 Instant: 26.5 (#259)

Knowledge GLM-5 leads

GLM-5: 52.3 (#64), GPT-5.5 Instant: 48.9 (#74)

Knowledge benchmarks
BenchmarkGLM-5GPT-5.5 Instant
GPQA Diamond87.8%82.5%
LMArena Expert14541409
Vectara Hallucination Rate10.1%—

Multimodal Not comparable

GLM-5: —, GPT-5.5 Instant: 40.0 (#52)

Multimodal benchmarks
BenchmarkGLM-5GPT-5.5 Instant
LMArena Vision—1250
LMArena Document—1403

Multilingual Too close to call

GLM-5: 53.7 (#58), GPT-5.5 Instant: 52.8 (#80)

Multilingual benchmarks
BenchmarkGLM-5GPT-5.5 Instant
LMArena Non-English14301417
LMArena Chinese15111456
LMArena French14551428
LMArena German14451411
LMArena Japanese14161408
LMArena Korean14231392
LMArena Russian14361431
LMArena Spanish14541429

Instruction Following GLM-5 leads

GLM-5: 75.2 (#67), GPT-5.5 Instant: 74.2 (#100)

Instruction Following benchmarks
BenchmarkGLM-5GPT-5.5 Instant
LMArena Instruction Following14281406

Long Context GLM-5 leads

GLM-5: 44.7 (#60), GPT-5.5 Instant: 43.4 (#96)

Long Context benchmarks
BenchmarkGLM-5GPT-5.5 Instant
LMArena Longer Query14461422
CL-bench18.7%—

Writing & Preference GLM-5 leads

GLM-5: 66.0 (#38), GPT-5.5 Instant: 61.8 (#85)

Writing & Preference benchmarks
BenchmarkGLM-5GPT-5.5 Instant
LMArena Text14461419
LMArena Creative Writing14391419
LMArena Multi-Turn14561433
EQ-Bench Creative Writing1601—

Frequently asked questions

Is GLM-5 better than GPT-5.5 Instant?

GLM-5 is the stronger model overall, scoring 46.1 to 42.7 on the Noometry Index.

Is GLM-5 or GPT-5.5 Instant better for coding?

GLM-5 scores higher on coding benchmarks: 49.0 versus 44.3 in the Noometry coding category.

How many benchmarks do GLM-5 and GPT-5.5 Instant share?

21 benchmarks have published results for both models. GLM-5 has 45 scored results on Noometry and GPT-5.5 Instant has 27.

Related comparisons

Go deeper