Model comparison

GLM-5.1 vs gpt-oss-120b

GLM-5.1 is the stronger model overall, scoring 47.8 to 36.3 on the Noometry Index. gpt-oss-120b costs 31× less per token, which makes it the better buy when GLM-5.1's lead doesn't matter for your workload.

Last verified . 29 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 29 benchmarks with published results for both. GLM-5.1 scores higher in 8 categories and gpt-oss-120b in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GLM-5.1 leads 66.9 to 46.5.
  • The biggest single-benchmark swing is APEX-Agents: 40.9% for GLM-5.1 and 4.4% for gpt-oss-120b.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $1.40 / $4.40 for GLM-5.1.
  • GLM-5.1 accepts more context: 200K tokens versus 131K.

Side by side

GLM-5.1 and gpt-oss-120b specifications
GLM-5.1gpt-oss-120b
ProviderZ.ai (Zhipu)OpenAI
Noometry Index47.836.3
Released2026-04-072025-08-05
WeightsOpenOpen
Context window200K131K
Max output131K41K
Input $ / M tokens$1.40$0.037
Output $ / M tokens$4.40$0.17
Results tracked4148

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.1 leads

GLM-5.1: 48.7 (#55), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkGLM-5.1gpt-oss-120b
SciCode43.8%36%
WeirdML57.1%48.2%
LMArena Coding14851380
ALE-Bench887.1575.62
SWE-bench Verified74.2%—
SWE-bench Verified (bash only)—26%
Aider Polyglot—41.8%
LMArena WebDev1508—
AlgoTune—1.41

Agentic & Tool Use GLM-5.1 leads

GLM-5.1: 24.9 (#113), gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1gpt-oss-120b
APEX-Agents40.9%4.4%
Vending-Bench 25,634-21.53
Terminal-Bench—18.7%
ExploitBench18.1%—
GBAEval0%—
METR Time Horizons—56.6%

Reasoning GLM-5.1 leads

GLM-5.1: 39.1 (#60), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkGLM-5.1gpt-oss-120b
SimpleBench55.1%22.1%
CritPt4.6%1.1%
Chess Puzzles19%20%
LMArena Hard Prompts14721364
Epoch Capabilities Index149.84139.93
Kagi LLM Benchmark—58.6%
NYT Connections (extended)77.7%—
Thematic Generalization69.8%—
Mystery Game Puzzles—2%
DTBench—76.3%
LMCA—22.1%
Surface Evolver Bench—25%

Math gpt-oss-120b leads

GLM-5.1: 49.7 (#60), gpt-oss-120b: 52.5 (#50)

Knowledge GLM-5.1 leads

GLM-5.1: 54.9 (#50), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkGLM-5.1gpt-oss-120b
GPQA Diamond89.9%75.8%
LMArena Expert14761356
SimpleQA Verified34%—
MMLU-Pro—79.5%
Confabulations—15.7%
Vectara Hallucination Rate—14.2%
GPQA (HELM)—68.4%

Multilingual GLM-5.1 leads

GLM-5.1: 55.0 (#36), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkGLM-5.1gpt-oss-120b
LMArena Non-English14471351
LMArena Chinese15151385
LMArena French14741369
LMArena German14651353
LMArena Japanese14341331
LMArena Korean14181282
LMArena Russian14541343
LMArena Spanish14691389

Instruction Following GLM-5.1 leads

GLM-5.1: 76.3 (#42), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkGLM-5.1gpt-oss-120b
LMArena Instruction Following14511318
IFEval—83.6%

Long Context GLM-5.1 leads

GLM-5.1: 44.9 (#53), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkGLM-5.1gpt-oss-120b
LMArena Longer Query14661319
Fiction.LiveBench—44.4%

Writing & Preference GLM-5.1 leads

GLM-5.1: 66.9 (#31), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkGLM-5.1gpt-oss-120b
LMArena Text14611365
LMArena Creative Writing14531275
EQ-Bench Creative Writing1592961
LMArena Multi-Turn14721340
Short-Story Creative Writing—77.1%
WildBench—84.5%

Frequently asked questions

Is GLM-5.1 better than gpt-oss-120b?

GLM-5.1 is the stronger model overall, scoring 47.8 to 36.3 on the Noometry Index. gpt-oss-120b costs 31× less per token, which makes it the better buy when GLM-5.1's lead doesn't matter for your workload.

Which is cheaper, GLM-5.1 or gpt-oss-120b?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; GLM-5.1 lists at $1.40 and $4.40.

Is GLM-5.1 or gpt-oss-120b better for coding?

GLM-5.1 scores higher on coding benchmarks: 48.7 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

GLM-5.1 does, with 200K tokens against 131K.

How many benchmarks do GLM-5.1 and gpt-oss-120b share?

29 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper