Model comparison

DeepSeek LLM 67B vs GLM-5V-Turbo

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 24.9 on the Noometry Index.

Last verified . 10 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

GLM-5V-Turbo Z.ai (Zhipu)

43.8

Rank #84 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and GLM-5V-Turbo in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-5V-Turbo leads 40.6 to 7.0.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and GLM-5V-Turbo specifications
DeepSeek LLM 67BGLM-5V-Turbo
ProviderDeepSeekZ.ai (Zhipu)
Noometry Index24.943.8
Released2023-11-292026-04-01
WeightsOpenProprietary
Context window—200K
Max output—131K
Input $ / M tokens—$1.20
Output $ / M tokens—$4
Results tracked1519

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5V-Turbo leads

DeepSeek LLM 67B: 31.9 (#278), GLM-5V-Turbo: 42.1 (#111)

Coding benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Coding10961466
LMArena WebDev—1401

Reasoning GLM-5V-Turbo leads

DeepSeek LLM 67B: 16.5 (#304), GLM-5V-Turbo: 29.7 (#89)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Hard Prompts10701443
Chess Puzzles0%—
Epoch Capabilities Index110.5—

Math GLM-5V-Turbo leads

DeepSeek LLM 67B: 8.7 (#324), GLM-5V-Turbo: 39.4 (#106)

Math benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Math11081441
OTIS Mock AIME 2024-20250.8%—
MATH Level 56.4%—

Knowledge GLM-5V-Turbo leads

DeepSeek LLM 67B: 7.0 (#313), GLM-5V-Turbo: 40.6 (#117)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
GPQA Diamond24.6%—
LMArena Expert—1452

Multimodal Not comparable

DeepSeek LLM 67B: —, GLM-5V-Turbo: 40.9 (#42)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Vision—1264
LMArena Document—1416

Multilingual GLM-5V-Turbo leads

DeepSeek LLM 67B: 29.4 (#267), GLM-5V-Turbo: 53.0 (#73)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Non-English10731420
LMArena Chinese11321488
LMArena French—1444
LMArena German—1423
LMArena Korean—1396
LMArena Russian—1431
LMArena Spanish—1450

Instruction Following GLM-5V-Turbo leads

DeepSeek LLM 67B: 55.4 (#277), GLM-5V-Turbo: 75.0 (#80)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Instruction Following10791423

Long Context GLM-5V-Turbo leads

DeepSeek LLM 67B: 33.1 (#265), GLM-5V-Turbo: 44.0 (#80)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Longer Query10921438

Writing & Preference GLM-5V-Turbo leads

DeepSeek LLM 67B: 31.6 (#282), GLM-5V-Turbo: 62.5 (#73)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BGLM-5V-Turbo
LMArena Text11051437
LMArena Creative Writing10671416
LMArena Multi-Turn10821432

Frequently asked questions

Is DeepSeek LLM 67B better than GLM-5V-Turbo?

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or GLM-5V-Turbo better for coding?

GLM-5V-Turbo scores higher on coding benchmarks: 42.1 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and GLM-5V-Turbo share?

10 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and GLM-5V-Turbo has 19.

Related comparisons

Go deeper