Model comparison

GLM-5V-Turbo vs Mercury 2

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 39.1 on the Noometry Index. Mercury 2 costs 5.1× less per token, which makes it the better buy when GLM-5V-Turbo's lead doesn't matter for your workload.

Last verified . 12 shared benchmarks.

GLM-5V-Turbo Z.ai (Zhipu)

43.8

Rank #84 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 12 benchmarks with published results for both. GLM-5V-Turbo scores higher in 7 categories and Mercury 2 in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GLM-5V-Turbo leads 62.5 to 53.8.
  • Mercury 2 is cheaper at $0.25 / $0.75 per million input/output tokens, against $1.20 / $4 for GLM-5V-Turbo.
  • GLM-5V-Turbo accepts more context: 200K tokens versus 128K.

Side by side

GLM-5V-Turbo and Mercury 2 specifications
GLM-5V-TurboMercury 2
ProviderZ.ai (Zhipu)Inception
Noometry Index43.839.1
Released2026-04-012026-02-20
WeightsProprietaryProprietary
Context window200K128K
Max output131K50K
Input $ / M tokens$1.20$0.25
Output $ / M tokens$4$0.75
Results tracked1917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5V-Turbo leads

GLM-5V-Turbo: 42.1 (#111), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena WebDev14011171
LMArena Coding14661391
SciCode—38.7%
WeirdML—43.2%
ALE-Bench—785.58

Reasoning GLM-5V-Turbo leads

GLM-5V-Turbo: 29.7 (#89), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Hard Prompts14431362
CritPt—0.8%

Math Not comparable

GLM-5V-Turbo: 39.4 (#106), Mercury 2: —

Math benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Math1441—

Knowledge GLM-5V-Turbo leads

GLM-5V-Turbo: 40.6 (#117), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Expert14521358
Vectara Hallucination Rate—12.3%

Multimodal Not comparable

GLM-5V-Turbo: 40.9 (#42), Mercury 2: —

Multimodal benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Vision1264—
LMArena Document1416—

Multilingual GLM-5V-Turbo leads

GLM-5V-Turbo: 53.0 (#73), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Non-English14201331
LMArena Chinese14881417
LMArena Russian14311304
LMArena French1444—
LMArena German1423—
LMArena Korean1396—
LMArena Spanish1450—

Instruction Following GLM-5V-Turbo leads

GLM-5V-Turbo: 75.0 (#80), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Instruction Following14231329

Long Context GLM-5V-Turbo leads

GLM-5V-Turbo: 44.0 (#80), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Longer Query14381330

Writing & Preference GLM-5V-Turbo leads

GLM-5V-Turbo: 62.5 (#73), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkGLM-5V-TurboMercury 2
LMArena Text14371355
LMArena Creative Writing14161289
LMArena Multi-Turn14321358

Frequently asked questions

Is GLM-5V-Turbo better than Mercury 2?

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 39.1 on the Noometry Index. Mercury 2 costs 5.1× less per token, which makes it the better buy when GLM-5V-Turbo's lead doesn't matter for your workload.

Which is cheaper, GLM-5V-Turbo or Mercury 2?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; GLM-5V-Turbo lists at $1.20 and $4.

Is GLM-5V-Turbo or Mercury 2 better for coding?

GLM-5V-Turbo scores higher on coding benchmarks: 42.1 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

GLM-5V-Turbo does, with 200K tokens against 128K.

How many benchmarks do GLM-5V-Turbo and Mercury 2 share?

12 benchmarks have published results for both models. GLM-5V-Turbo has 19 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper