Model comparison

Claude Opus 4.1 vs GLM-5V-Turbo

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 41.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

GLM-5V-Turbo Z.ai (Zhipu)

43.8

Rank #84 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude Opus 4.1 scores higher in 5 categories and GLM-5V-Turbo in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5V-Turbo leads 39.4 to 22.3.
  • GLM-5V-Turbo is cheaper at $1.20 / $4 per million input/output tokens, against $15 / $75 for Claude Opus 4.1.

Side by side

Claude Opus 4.1 and GLM-5V-Turbo specifications
Claude Opus 4.1GLM-5V-Turbo
ProviderAnthropicZ.ai (Zhipu)
Noometry Index41.043.8
Released2025-08-052026-04-01
WeightsProprietaryProprietary
Context window200K200K
Max output32K131K
Input $ / M tokens$15$1.20
Output $ / M tokens$75$4
Results tracked4819

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.1 leads

Claude Opus 4.1: 44.4 (#73), GLM-5V-Turbo: 42.1 (#111)

Coding benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena WebDev13901401
LMArena Coding14791466
SWE-bench Verified73.3%—
WeirdML45.9%—
ALE-Bench674.77—
AlgoTune1.34—

Agentic & Tool Use Not comparable

Claude Opus 4.1: 35.0 (#41), GLM-5V-Turbo: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
Terminal-Bench38%—
GDPval43.6%—
Cybench42%—
DeepResearch Bench48.3%—
LMArena Search1148—
METR Time Horizons66.8%—

Reasoning Claude Opus 4.1 leads

Claude Opus 4.1: 32.2 (#76), GLM-5V-Turbo: 29.7 (#89)

Reasoning benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Hard Prompts14431443
SimpleBench60%—
Chess Puzzles7%—
EnigmaEval7.2%—
EBR-Bench7.9%—
Mystery Game Puzzles21%—
DTBench80%—
LMCA37.1%—
Epoch Capabilities Index144.12—
ForecastBench62—

Math GLM-5V-Turbo leads

Claude Opus 4.1: 22.3 (#277), GLM-5V-Turbo: 39.4 (#106)

Math benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Math14311441
FrontierMath (Tiers 1-3)12.6%—
FrontierMath Tier 42.4%—
OTIS Mock AIME 2024-202568.9%—
FrontierMath (Feb 2025 set)7.2%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Claude Opus 4.1 leads

Claude Opus 4.1: 42.0 (#101), GLM-5V-Turbo: 40.6 (#117)

Knowledge benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Expert14391452
GPQA Diamond77.3%—
Humanity's Last Exam11.5%—
Confabulations17.1%—
Vectara Hallucination Rate11.8%—

Multimodal GLM-5V-Turbo leads

Claude Opus 4.1: 26.8 (#119), GLM-5V-Turbo: 40.9 (#42)

Multimodal benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Vision—1264
VPCT35%—
LMArena Document—1416

Multilingual GLM-5V-Turbo leads

Claude Opus 4.1: 52.0 (#95), GLM-5V-Turbo: 53.0 (#73)

Multilingual benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Non-English14051420
LMArena Chinese14271488
LMArena French14311444
LMArena German14131423
LMArena Korean13801396
LMArena Russian14221431
LMArena Spanish14481450
LMArena Japanese1378—

Instruction Following Too close to call

Claude Opus 4.1: 75.6 (#58), GLM-5V-Turbo: 75.0 (#80)

Instruction Following benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Instruction Following14351423

Long Context Too close to call

Claude Opus 4.1: 44.5 (#63), GLM-5V-Turbo: 44.0 (#80)

Long Context benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Longer Query14551438

Writing & Preference Too close to call

Claude Opus 4.1: 62.4 (#74), GLM-5V-Turbo: 62.5 (#73)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1GLM-5V-Turbo
LMArena Text14191437
LMArena Creative Writing14121416
LMArena Multi-Turn14441432
Short-Story Creative Writing84.7%—

Frequently asked questions

Is Claude Opus 4.1 better than GLM-5V-Turbo?

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4.1 or GLM-5V-Turbo?

GLM-5V-Turbo is cheaper. It lists at $1.20 per million input tokens and $4 per million output tokens; Claude Opus 4.1 lists at $15 and $75.

Is Claude Opus 4.1 or GLM-5V-Turbo better for coding?

Claude Opus 4.1 scores higher on coding benchmarks: 44.4 versus 42.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 200K tokens.

How many benchmarks do Claude Opus 4.1 and GLM-5V-Turbo share?

17 benchmarks have published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and GLM-5V-Turbo has 19.

Related comparisons

Go deeper