Model comparison

GLM-4.7-Flash vs Mercury

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 37.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. GLM-4.7-Flash scores higher in 6 categories and Mercury in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GLM-4.7-Flash leads 46.5 to 41.6.
  • GLM-4.7-Flash has downloadable open weights; the other is API-only.

Side by side

GLM-4.7-Flash and Mercury specifications
GLM-4.7-FlashMercury
ProviderZ.ai (Zhipu)Inception
Noometry Index38.837.6
Released2026-01-19—
WeightsOpenProprietary
Context window200K—
Max output131K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.40—
Results tracked219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

GLM-4.7-Flash: 40.6 (#135), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkGLM-4.7-FlashMercury
LMArena Coding13831322

Reasoning GLM-4.7-Flash leads

GLM-4.7-Flash: 20.9 (#229), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashMercury
LMArena Hard Prompts13561285
Kagi LLM Benchmark—21.6%
Chess Puzzles0%—

Math Not comparable

GLM-4.7-Flash: 36.1 (#173), Mercury: —

Math benchmarks
BenchmarkGLM-4.7-FlashMercury
OTIS Mock AIME 2024-202558.3%—
LMArena Math1355—

Knowledge Not comparable

GLM-4.7-Flash: 35.5 (#184), Mercury: —

Knowledge benchmarks
BenchmarkGLM-4.7-FlashMercury
GPQA Diamond60.5%—
Vectara Hallucination Rate9.3%—
LMArena Expert1357—

Multilingual GLM-4.7-Flash leads

GLM-4.7-Flash: 46.5 (#158), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashMercury
LMArena Non-English13301260
LMArena Chinese1403—
LMArena French1332—
LMArena German1337—
LMArena Korean1283—
LMArena Russian1332—
LMArena Spanish1350—

Instruction Following GLM-4.7-Flash leads

GLM-4.7-Flash: 70.1 (#167), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashMercury
LMArena Instruction Following13271239

Long Context GLM-4.7-Flash leads

GLM-4.7-Flash: 40.9 (#148), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkGLM-4.7-FlashMercury
LMArena Longer Query13451266

Writing & Preference GLM-4.7-Flash leads

GLM-4.7-Flash: 47.4 (#210), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashMercury
LMArena Text13511282
LMArena Creative Writing12971191
LMArena Multi-Turn13421282
EQ-Bench Creative Writing1125—

Frequently asked questions

Is GLM-4.7-Flash better than Mercury?

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 37.6 on the Noometry Index.

Is GLM-4.7-Flash or Mercury better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 38.7 in the Noometry coding category.

How many benchmarks do GLM-4.7-Flash and Mercury share?

8 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper