Model comparison

Command R vs DeepSeek LLM 67B

Command R is the stronger model overall, scoring 31.4 to 24.9 on the Noometry Index.

Last verified . 10 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Command R scores higher in 6 categories and DeepSeek LLM 67B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command R leads 31.0 to 7.0.

Side by side

Command R and DeepSeek LLM 67B specifications
Command RDeepSeek LLM 67B
ProviderCohereDeepSeek
Noometry Index31.424.9
Released2024-08-302023-11-29
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek LLM 67B leads

Command R: 29.3 (#306), DeepSeek LLM 67B: 31.9 (#278)

Coding benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Coding11691096
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Reasoning DeepSeek LLM 67B leads

Command R: 13.8 (#331), DeepSeek LLM 67B: 16.5 (#304)

Reasoning benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Hard Prompts11641070
Chess Puzzles—0%
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
Epoch Capabilities Index—110.5
LiveBench27.5%—

Math Command R leads

Command R: 28.0 (#246), DeepSeek LLM 67B: 8.7 (#324)

Math benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Math11551108
OTIS Mock AIME 2024-2025—0.8%
LiveBench Math19.4%—
MATH Level 5—6.4%

Knowledge Command R leads

Command R: 31.0 (#221), DeepSeek LLM 67B: 7.0 (#313)

Knowledge benchmarks
BenchmarkCommand RDeepSeek LLM 67B
GPQA Diamond—24.6%
LMArena Expert1138—
MMLU65.2%—

Multilingual Command R leads

Command R: 35.7 (#245), DeepSeek LLM 67B: 29.4 (#267)

Multilingual benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Non-English11741073
LMArena Chinese11821132
LMArena French1162—
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Russian1174—
LMArena Spanish1151—

Instruction Following Command R leads

Command R: 58.1 (#261), DeepSeek LLM 67B: 55.4 (#277)

Instruction Following benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Instruction Following11671079
LiveBench Instruction Following55.6%—

Long Context Command R leads

Command R: 36.3 (#231), DeepSeek LLM 67B: 33.1 (#265)

Long Context benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Longer Query11981092

Writing & Preference Command R leads

Command R: 38.2 (#254), DeepSeek LLM 67B: 31.6 (#282)

Writing & Preference benchmarks
BenchmarkCommand RDeepSeek LLM 67B
LMArena Text11871105
LMArena Creative Writing11701067
LMArena Multi-Turn11631082
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than DeepSeek LLM 67B?

Command R is the stronger model overall, scoring 31.4 to 24.9 on the Noometry Index.

Is Command R or DeepSeek LLM 67B better for coding?

DeepSeek LLM 67B scores higher on coding benchmarks: 31.9 versus 29.3 in the Noometry coding category.

How many benchmarks do Command R and DeepSeek LLM 67B share?

10 benchmarks have published results for both models. Command R has 29 scored results on Noometry and DeepSeek LLM 67B has 15.

Related comparisons

Go deeper