Model comparison

DeepSeek-V3.1 vs GLM-4.7-Flash

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 38.8 on the Noometry Index. GLM-4.7-Flash costs 2.9× less per token, which makes it the better buy when DeepSeek-V3.1's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 6 categories and GLM-4.7-Flash in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.1 leads 60.3 to 47.4.
  • GLM-4.7-Flash is cheaper at $0.06 / $0.40 per million input/output tokens, against $0.25 / $0.95 for DeepSeek-V3.1.
  • GLM-4.7-Flash accepts more context: 200K tokens versus 164K.

Side by side

DeepSeek-V3.1 and GLM-4.7-Flash specifications
DeepSeek-V3.1GLM-4.7-Flash
ProviderDeepSeekZ.ai (Zhipu)
Noometry Index42.838.8
Released2025-08-212026-01-19
WeightsOpenOpen
Context window164K200K
Max output8K131K
Input $ / M tokens$0.25$0.06
Output $ / M tokens$0.95$0.40
Results tracked2721

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V3.1: 40.3 (#144), GLM-4.7-Flash: 40.6 (#135)

Coding benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Coding14171383
WeirdML38.4%—

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), GLM-4.7-Flash: 20.9 (#229)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Hard Prompts14171356
SimpleBench40%—
Kagi LLM Benchmark53.2%—
Chess Puzzles—0%
DTBench82.7%—
LMCA24.3%—
Epoch Capabilities Index139.92—
ForecastBench58—

Math DeepSeek-V3.1 leads

DeepSeek-V3.1: 38.9 (#122), GLM-4.7-Flash: 36.1 (#173)

Math benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Math14201355
OTIS Mock AIME 2024-2025—58.3%

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), GLM-4.7-Flash: 35.5 (#184)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
Vectara Hallucination Rate5.5%9.3%
LMArena Expert14051357
GPQA Diamond—60.5%

Multilingual DeepSeek-V3.1 leads

DeepSeek-V3.1: 51.6 (#106), GLM-4.7-Flash: 46.5 (#158)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Non-English14001330
LMArena Chinese14691403
LMArena French14471332
LMArena German14111337
LMArena Korean13371283
LMArena Russian14051332
LMArena Spanish14311350
LMArena Japanese1378—

Instruction Following DeepSeek-V3.1 leads

DeepSeek-V3.1: 73.9 (#110), GLM-4.7-Flash: 70.1 (#167)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Instruction Following14001327

Long Context GLM-4.7-Flash leads

DeepSeek-V3.1: 36.3 (#232), GLM-4.7-Flash: 40.9 (#148)

Long Context benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Longer Query14221345
Fiction.LiveBench52.8%—

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), GLM-4.7-Flash: 47.4 (#210)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1GLM-4.7-Flash
LMArena Text14201351
LMArena Creative Writing14011297
EQ-Bench Creative Writing14361125
LMArena Multi-Turn14081342

Frequently asked questions

Is DeepSeek-V3.1 better than GLM-4.7-Flash?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 38.8 on the Noometry Index. GLM-4.7-Flash costs 2.9× less per token, which makes it the better buy when DeepSeek-V3.1's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3.1 or GLM-4.7-Flash?

GLM-4.7-Flash is cheaper. It lists at $0.06 per million input tokens and $0.40 per million output tokens; DeepSeek-V3.1 lists at $0.25 and $0.95.

Is DeepSeek-V3.1 or GLM-4.7-Flash better for coding?

They score almost the same on coding (40.3 vs 40.6); test both on your own repository before choosing.

Which has the bigger context window?

GLM-4.7-Flash does, with 200K tokens against 164K.

How many benchmarks do DeepSeek-V3.1 and GLM-4.7-Flash share?

18 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and GLM-4.7-Flash has 21.

Related comparisons

Go deeper