Model comparison

Command A vs DeepSeek-V3.2-Speciale

DeepSeek-V3.2-Speciale is the stronger model overall, scoring 39.7 to 36.5 on the Noometry Index.

Last verified . 1 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

DeepSeek-V3.2-Speciale DeepSeek

39.7

Rank #162 Reported

Summary

  • They share 1 benchmark with published results for both. Command A scores higher in 1 category and DeepSeek-V3.2-Speciale in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.2-Speciale leads 32.9 to 18.3.
  • DeepSeek-V3.2-Speciale is cheaper at $0.58 / $1.68 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A accepts more context: 256K tokens versus 128K.

Side by side

Command A and DeepSeek-V3.2-Speciale specifications
Command ADeepSeek-V3.2-Speciale
ProviderCohereDeepSeek
Noometry Index36.539.7
Released2025-03-132025-12-01
WeightsOpenOpen
Context window256K128K
Max output8K128K
Input $ / M tokens$2.50$0.58
Output $ / M tokens$10$1.68
Results tracked243

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Speciale leads

Command A: 27.2 (#322), DeepSeek-V3.2-Speciale: 40.4 (#140)

Coding benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
Aider Polyglot12%—
WeirdML—46.7%
LMArena Coding1330—

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), DeepSeek-V3.2-Speciale: —

Agentic & Tool Use benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
Berkeley Function Calling Leaderboard57.1%—

Reasoning DeepSeek-V3.2-Speciale leads

Command A: 18.3 (#283), DeepSeek-V3.2-Speciale: 32.9 (#73)

Reasoning benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
SimpleBench—52.6%
Kagi LLM Benchmark28.8%—
LMArena Hard Prompts1326—
DTBench61.3%—
LMCA10.3%—

Math Not comparable

Command A: 36.2 (#171), DeepSeek-V3.2-Speciale: —

Math benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
LMArena Math1300—

Knowledge Not comparable

Command A: 37.1 (#159), DeepSeek-V3.2-Speciale: —

Knowledge benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
Vectara Hallucination Rate9.3%—
LMArena Expert1295—

Multilingual Not comparable

Command A: 45.3 (#170), DeepSeek-V3.2-Speciale: —

Multilingual benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
LMArena Non-English1313—
LMArena Chinese1327—
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Russian1314—
LMArena Spanish1347—

Instruction Following Not comparable

Command A: 69.1 (#177), DeepSeek-V3.2-Speciale: —

Instruction Following benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
LMArena Instruction Following1309—

Long Context Not comparable

Command A: 40.6 (#151), DeepSeek-V3.2-Speciale: —

Long Context benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
LMArena Longer Query1334—

Writing & Preference Command A leads

Command A: 47.6 (#208), DeepSeek-V3.2-Speciale: 46.0 (#222)

Writing & Preference benchmarks
BenchmarkCommand ADeepSeek-V3.2-Speciale
EQ-Bench Creative Writing11451276
LMArena Text1331—
LMArena Creative Writing1319—
LMArena Multi-Turn1339—

Frequently asked questions

Is Command A better than DeepSeek-V3.2-Speciale?

DeepSeek-V3.2-Speciale is the stronger model overall, scoring 39.7 to 36.5 on the Noometry Index.

Which is cheaper, Command A or DeepSeek-V3.2-Speciale?

DeepSeek-V3.2-Speciale is cheaper. It lists at $0.58 per million input tokens and $1.68 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or DeepSeek-V3.2-Speciale better for coding?

DeepSeek-V3.2-Speciale scores higher on coding benchmarks: 40.4 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Command A does, with 256K tokens against 128K.

How many benchmarks do Command A and DeepSeek-V3.2-Speciale share?

1 benchmark has published results for both models. Command A has 24 scored results on Noometry and DeepSeek-V3.2-Speciale has 3.

Related comparisons

Go deeper