Model comparison

Claude 3 Opus vs GLM-5.1

GLM-5.1 is the stronger model overall, scoring 47.8 to 29.5 on the Noometry Index.

Last verified . 24 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Claude 3 Opus scores higher in 0 categories and GLM-5.1 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5.1 leads 49.7 to 14.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.7% for Claude 3 Opus and 93.3% for GLM-5.1.
  • GLM-5.1 has downloadable open weights; the other is API-only.

Side by side

Claude 3 Opus and GLM-5.1 specifications
Claude 3 OpusGLM-5.1
ProviderAnthropicZ.ai (Zhipu)
Noometry Index29.547.8
Released2024-02-292026-04-07
WeightsProprietaryOpen
Context window—200K
Max output—131K
Input $ / M tokens—$1.40
Output $ / M tokens—$4.40
Results tracked4641

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.1 leads

Claude 3 Opus: 32.9 (#267), GLM-5.1: 48.7 (#55)

Coding benchmarks
BenchmarkClaude 3 OpusGLM-5.1
WeirdML19.2%57.1%
LMArena Coding12641485
SWE-bench Verified—74.2%
LMArena WebDev—1508
SciCode—43.8%
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
BigCodeBench Complete57.4%—
ALE-Bench—887.1
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use Too close to call

Claude 3 Opus: 24.6 (#116), GLM-5.1: 24.9 (#113)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGLM-5.1
APEX-Agents—40.9%
Cybench10%—
ExploitBench—18.1%
GBAEval—0%
METR Time Horizons29.5%—
Vending-Bench 2—5,634

Reasoning GLM-5.1 leads

Claude 3 Opus: 14.6 (#324), GLM-5.1: 39.1 (#60)

Reasoning benchmarks
BenchmarkClaude 3 OpusGLM-5.1
SimpleBench23.5%55.1%
Chess Puzzles5%19%
LMArena Hard Prompts12451472
Epoch Capabilities Index126.91149.84
NYT Connections (extended)—77.7%
CritPt—4.6%
EnigmaEval0.8%—
Thematic Generalization—69.8%
LiveBench Reasoning40.6%—
DTBench61.6%—
LiveBench Data Analysis57.9%—
LMCA17%—
ForecastBench58.4—
LiveBench49.2%—
WinoGrande88.5%—

Math GLM-5.1 leads

Claude 3 Opus: 14.8 (#299), GLM-5.1: 49.7 (#60)

Knowledge GLM-5.1 leads

Claude 3 Opus: 24.5 (#267), GLM-5.1: 54.9 (#50)

Knowledge benchmarks
BenchmarkClaude 3 OpusGLM-5.1
GPQA Diamond47.2%89.9%
SimpleQA Verified12.6%34%
LMArena Expert12231476
Confabulations22.7%—
MMLU84.6%—

Multimodal Not comparable

Claude 3 Opus: 27.1 (#116), GLM-5.1: —

Multimodal benchmarks
BenchmarkClaude 3 OpusGLM-5.1
LMArena Vision1023—

Multilingual GLM-5.1 leads

Claude 3 Opus: 41.4 (#207), GLM-5.1: 55.0 (#36)

Multilingual benchmarks
BenchmarkClaude 3 OpusGLM-5.1
LMArena Non-English12581447
LMArena Chinese12481515
LMArena French12751474
LMArena German12581465
LMArena Japanese12041434
LMArena Korean11871418
LMArena Russian12801454
LMArena Spanish12461469

Instruction Following GLM-5.1 leads

Claude 3 Opus: 64.1 (#228), GLM-5.1: 76.3 (#42)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGLM-5.1
LMArena Instruction Following12481451
LiveBench Instruction Following63.9%—

Long Context GLM-5.1 leads

Claude 3 Opus: 38.2 (#202), GLM-5.1: 44.9 (#53)

Long Context benchmarks
BenchmarkClaude 3 OpusGLM-5.1
LMArena Longer Query12591466

Writing & Preference GLM-5.1 leads

Claude 3 Opus: 47.2 (#213), GLM-5.1: 66.9 (#31)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGLM-5.1
LMArena Text12621461
LMArena Creative Writing12351453
LMArena Multi-Turn12751472
EQ-Bench Creative Writing—1592
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than GLM-5.1?

GLM-5.1 is the stronger model overall, scoring 47.8 to 29.5 on the Noometry Index.

Is Claude 3 Opus or GLM-5.1 better for coding?

GLM-5.1 scores higher on coding benchmarks: 48.7 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and GLM-5.1 share?

24 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and GLM-5.1 has 41.

Related comparisons

Go deeper