Model comparison

GLM-4.6 vs GPT-5.5 Instant

GPT-5.5 Instant is the stronger model overall, scoring 42.7 to 41.4 on the Noometry Index.

Last verified . 19 shared benchmarks.

GLM-4.6 Z.ai (Zhipu)

41.4

Rank #135 Confirmed

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Summary

  • They share 19 benchmarks with published results for both. GLM-4.6 scores higher in 4 categories and GPT-5.5 Instant in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-4.6 leads 39.1 to 26.5.
  • The biggest single-benchmark swing is SciCode: 38.4% for GLM-4.6 and 48.6% for GPT-5.5 Instant.
  • GLM-4.6 has downloadable open weights; the other is API-only.

Side by side

GLM-4.6 and GPT-5.5 Instant specifications
GLM-4.6GPT-5.5 Instant
ProviderZ.ai (Zhipu)OpenAI
Noometry Index41.442.7
Released2025-09-302026-05-05
WeightsOpenProprietary
Context window205K—
Max output131K—
Input $ / M tokens$0.60—
Output $ / M tokens$2.20—
Results tracked2927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 Instant leads

GLM-4.6: 40.1 (#148), GPT-5.5 Instant: 44.3 (#74)

Coding benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
SciCode38.4%48.6%
LMArena Coding14491433
SWE-bench Verified (bash only)55.4%—
LMArena WebDev1340—
ALE-Bench340.82—

Agentic & Tool Use Not comparable

GLM-4.6: 32.3 (#66), GPT-5.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
Terminal-Bench24.5%—
Berkeley Function Calling Leaderboard72.4%—

Reasoning GPT-5.5 Instant leads

GLM-4.6: 23.7 (#172), GPT-5.5 Instant: 24.9 (#155)

Reasoning benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
CritPt1.1%0%
LMArena Hard Prompts14401426
Kagi LLM Benchmark47.4%—
Chess Puzzles—12%
Epoch Capabilities Index—142.52

Math GLM-4.6 leads

GLM-4.6: 39.1 (#111), GPT-5.5 Instant: 26.5 (#259)

Math benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Math14321420
FrontierMath (Tiers 1-3)—26.3%
FrontierMath Tier 4—2.4%
OTIS Mock AIME 2024-2025—68.1%
FrontierMath (Feb 2025 set)3.8%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge GPT-5.5 Instant leads

GLM-4.6: 40.2 (#124), GPT-5.5 Instant: 48.9 (#74)

Knowledge benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Expert14311409
GPQA Diamond—82.5%
Vectara Hallucination Rate9.5%—

Multimodal Not comparable

GLM-4.6: —, GPT-5.5 Instant: 40.0 (#52)

Multimodal benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Vision—1250
LMArena Document—1403

Multilingual Too close to call

GLM-4.6: 53.5 (#66), GPT-5.5 Instant: 52.8 (#80)

Multilingual benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Non-English14261417
LMArena Chinese14991456
LMArena French14591428
LMArena German14471411
LMArena Japanese13931408
LMArena Korean14001392
LMArena Russian14191431
LMArena Spanish14361429

Instruction Following Too close to call

GLM-4.6: 74.3 (#98), GPT-5.5 Instant: 74.2 (#100)

Instruction Following benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Instruction Following14101406

Long Context Too close to call

GLM-4.6: 43.4 (#94), GPT-5.5 Instant: 43.4 (#96)

Long Context benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Longer Query14221422

Writing & Preference Too close to call

GLM-4.6: 61.1 (#90), GPT-5.5 Instant: 61.8 (#85)

Writing & Preference benchmarks
BenchmarkGLM-4.6GPT-5.5 Instant
LMArena Text14401419
LMArena Creative Writing14111419
LMArena Multi-Turn14271433
EQ-Bench Creative Writing1411—

Frequently asked questions

Is GLM-4.6 better than GPT-5.5 Instant?

GPT-5.5 Instant is the stronger model overall, scoring 42.7 to 41.4 on the Noometry Index.

Is GLM-4.6 or GPT-5.5 Instant better for coding?

GPT-5.5 Instant scores higher on coding benchmarks: 44.3 versus 40.1 in the Noometry coding category.

How many benchmarks do GLM-4.6 and GPT-5.5 Instant share?

19 benchmarks have published results for both models. GLM-4.6 has 29 scored results on Noometry and GPT-5.5 Instant has 27.

Related comparisons

Go deeper