Model comparison

DeepSeek LLM 67B vs GLM-4.7-Flash

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 24.9 on the Noometry Index.

Last verified . 13 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Summary

  • They share 13 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and GLM-4.7-Flash in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-4.7-Flash leads 35.5 to 7.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 0.8% for DeepSeek LLM 67B and 58.3% for GLM-4.7-Flash.

Side by side

DeepSeek LLM 67B and GLM-4.7-Flash specifications
DeepSeek LLM 67BGLM-4.7-Flash
ProviderDeepSeekZ.ai (Zhipu)
Noometry Index24.938.8
Released2023-11-292026-01-19
WeightsOpenOpen
Context window—200K
Max output—131K
Input $ / M tokens—$0.06
Output $ / M tokens—$0.40
Results tracked1521

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

DeepSeek LLM 67B: 31.9 (#278), GLM-4.7-Flash: 40.6 (#135)

Coding benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
LMArena Coding10961383

Reasoning GLM-4.7-Flash leads

DeepSeek LLM 67B: 16.5 (#304), GLM-4.7-Flash: 20.9 (#229)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
Chess Puzzles0%0%
LMArena Hard Prompts10701356
Epoch Capabilities Index110.5—

Math GLM-4.7-Flash leads

DeepSeek LLM 67B: 8.7 (#324), GLM-4.7-Flash: 36.1 (#173)

Math benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
OTIS Mock AIME 2024-20250.8%58.3%
LMArena Math11081355
MATH Level 56.4%—

Knowledge GLM-4.7-Flash leads

DeepSeek LLM 67B: 7.0 (#313), GLM-4.7-Flash: 35.5 (#184)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
GPQA Diamond24.6%60.5%
Vectara Hallucination Rate—9.3%
LMArena Expert—1357

Multilingual GLM-4.7-Flash leads

DeepSeek LLM 67B: 29.4 (#267), GLM-4.7-Flash: 46.5 (#158)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
LMArena Non-English10731330
LMArena Chinese11321403
LMArena French—1332
LMArena German—1337
LMArena Korean—1283
LMArena Russian—1332
LMArena Spanish—1350

Instruction Following GLM-4.7-Flash leads

DeepSeek LLM 67B: 55.4 (#277), GLM-4.7-Flash: 70.1 (#167)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
LMArena Instruction Following10791327

Long Context GLM-4.7-Flash leads

DeepSeek LLM 67B: 33.1 (#265), GLM-4.7-Flash: 40.9 (#148)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
LMArena Longer Query10921345

Writing & Preference GLM-4.7-Flash leads

DeepSeek LLM 67B: 31.6 (#282), GLM-4.7-Flash: 47.4 (#210)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.7-Flash
LMArena Text11051351
LMArena Creative Writing10671297
LMArena Multi-Turn10821342
EQ-Bench Creative Writing—1125

Frequently asked questions

Is DeepSeek LLM 67B better than GLM-4.7-Flash?

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or GLM-4.7-Flash better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and GLM-4.7-Flash share?

13 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and GLM-4.7-Flash has 21.

Related comparisons

Go deeper