Model comparison

DeepSeek LLM 67B vs GLM-4.5V

GLM-4.5V is the stronger model overall, scoring 39.8 to 24.9 on the Noometry Index.

Last verified . 10 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

GLM-4.5V Z.ai (Zhipu)

39.8

Rank #158 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and GLM-4.5V in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-4.5V leads 37.5 to 7.0.

Side by side

DeepSeek LLM 67B and GLM-4.5V specifications
DeepSeek LLM 67BGLM-4.5V
ProviderDeepSeekZ.ai (Zhipu)
Noometry Index24.939.8
Released2023-11-292025-08-11
WeightsOpenOpen
Context window—64K
Max output—16K
Input $ / M tokens—$0.60
Output $ / M tokens—$1.80
Results tracked1515

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.5V leads

DeepSeek LLM 67B: 31.9 (#278), GLM-4.5V: 39.5 (#155)

Coding benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Coding10961347

Reasoning GLM-4.5V leads

DeepSeek LLM 67B: 16.5 (#304), GLM-4.5V: 27.4 (#119)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Hard Prompts10701334
Kagi LLM Benchmark—59.8%
Chess Puzzles0%—
Epoch Capabilities Index110.5—

Math GLM-4.5V leads

DeepSeek LLM 67B: 8.7 (#324), GLM-4.5V: 37.4 (#159)

Math benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Math11081354
OTIS Mock AIME 2024-20250.8%—
MATH Level 56.4%—

Knowledge GLM-4.5V leads

DeepSeek LLM 67B: 7.0 (#313), GLM-4.5V: 37.5 (#156)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
GPQA Diamond24.6%—
LMArena Expert—1353

Multimodal Not comparable

DeepSeek LLM 67B: —, GLM-4.5V: 34.3 (#92)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Vision—1154

Multilingual GLM-4.5V leads

DeepSeek LLM 67B: 29.4 (#267), GLM-4.5V: 44.6 (#177)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Non-English10731303
LMArena Chinese11321337
LMArena Russian—1298
LMArena Spanish—1336

Instruction Following GLM-4.5V leads

DeepSeek LLM 67B: 55.4 (#277), GLM-4.5V: 69.2 (#175)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Instruction Following10791311

Long Context GLM-4.5V leads

DeepSeek LLM 67B: 33.1 (#265), GLM-4.5V: 39.6 (#171)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Longer Query10921304

Writing & Preference GLM-4.5V leads

DeepSeek LLM 67B: 31.6 (#282), GLM-4.5V: 52.5 (#170)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BGLM-4.5V
LMArena Text11051333
LMArena Creative Writing10671295
LMArena Multi-Turn10821332

Frequently asked questions

Is DeepSeek LLM 67B better than GLM-4.5V?

GLM-4.5V is the stronger model overall, scoring 39.8 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or GLM-4.5V better for coding?

GLM-4.5V scores higher on coding benchmarks: 39.5 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and GLM-4.5V share?

10 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and GLM-4.5V has 15.

Related comparisons

Go deeper