Model comparison

Command A vs DeepSeek-V3.2-Exp

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 36.5 on the Noometry Index.

Last verified . 24 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Command A scores higher in 1 category and DeepSeek-V3.2-Exp in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where DeepSeek-V3.2-Exp leads 46.5 to 27.2.
  • The biggest single-benchmark swing is Aider Polyglot: 12% for Command A and 74.2% for DeepSeek-V3.2-Exp.
  • DeepSeek-V3.2-Exp is cheaper at $0.26 / $0.38 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A accepts more context: 256K tokens versus 164K.

Side by side

Command A and DeepSeek-V3.2-Exp specifications
Command ADeepSeek-V3.2-Exp
ProviderCohereDeepSeek
Noometry Index36.544.3
Released2025-03-132025-09-29
WeightsOpenOpen
Context window256K164K
Max output8K66K
Input $ / M tokens$2.50$0.26
Output $ / M tokens$10$0.38
Results tracked2449

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

Command A: 27.2 (#322), DeepSeek-V3.2-Exp: 46.5 (#65)

Coding benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
Aider Polyglot12%74.2%
LMArena Coding13301454
SWE-bench Verified (bash only)—70%
LMArena WebDev—1362
SWE-bench Multilingual—59%
SciCode—38.9%
WeirdML—39.5%

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), DeepSeek-V3.2-Exp: 32.7 (#59)

Agentic & Tool Use benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
Berkeley Function Calling Leaderboard57.1%56.7%
Terminal-Bench—39.6%
APEX-Agents—21.3%
TheAgentCompany—42.9%
Vending-Bench 2—1,034

Reasoning DeepSeek-V3.2-Exp leads

Command A: 18.3 (#283), DeepSeek-V3.2-Exp: 22.1 (#208)

Reasoning benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
Kagi LLM Benchmark28.8%52.2%
LMArena Hard Prompts13261434
DTBench61.3%87.7%
LMCA10.3%29.1%
ARC-AGI-2—4%
NYT Connections (extended)—36.7%
ARC-AGI-1—57%
CritPt—2.9%
Chess Puzzles—14%
Thematic Generalization—65%
Epoch Capabilities Index—146.27

Math DeepSeek-V3.2-Exp leads

Command A: 36.2 (#171), DeepSeek-V3.2-Exp: 41.7 (#87)

Math benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
LMArena Math13001435
MathArena Final-Answer Competitions—57.7%
OTIS Mock AIME 2024-2025—87.8%
ProofBench—8%
FrontierMath (Feb 2025 set)—22.1%
FrontierMath Tier 4 (v1)—2.1%

Knowledge DeepSeek-V3.2-Exp leads

Command A: 37.1 (#159), DeepSeek-V3.2-Exp: 51.7 (#66)

Knowledge benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
Vectara Hallucination Rate9.3%5.3%
LMArena Expert12951436
GPQA Diamond—83.4%

Multilingual DeepSeek-V3.2-Exp leads

Command A: 45.3 (#170), DeepSeek-V3.2-Exp: 52.2 (#90)

Multilingual benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
LMArena Non-English13131409
LMArena Chinese13271461
LMArena French13511433
LMArena German13411440
LMArena Japanese12851374
LMArena Korean12851371
LMArena Russian13141424
LMArena Spanish13471440

Instruction Following DeepSeek-V3.2-Exp leads

Command A: 69.1 (#177), DeepSeek-V3.2-Exp: 74.5 (#93)

Instruction Following benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
LMArena Instruction Following13091413

Long Context DeepSeek-V3.2-Exp leads

Command A: 40.6 (#151), DeepSeek-V3.2-Exp: 47.6 (#16)

Long Context benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
LMArena Longer Query13341428
Fiction.LiveBench—83.3%
CL-bench—13.2%
CL-bench Life—9.5%

Writing & Preference DeepSeek-V3.2-Exp leads

Command A: 47.6 (#208), DeepSeek-V3.2-Exp: 62.4 (#77)

Writing & Preference benchmarks
BenchmarkCommand ADeepSeek-V3.2-Exp
LMArena Text13311425
LMArena Creative Writing13191403
EQ-Bench Creative Writing11451515
LMArena Multi-Turn13391427

Frequently asked questions

Is Command A better than DeepSeek-V3.2-Exp?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 36.5 on the Noometry Index.

Which is cheaper, Command A or DeepSeek-V3.2-Exp?

DeepSeek-V3.2-Exp is cheaper. It lists at $0.26 per million input tokens and $0.38 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or DeepSeek-V3.2-Exp better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Command A does, with 256K tokens against 164K.

How many benchmarks do Command A and DeepSeek-V3.2-Exp share?

24 benchmarks have published results for both models. Command A has 24 scored results on Noometry and DeepSeek-V3.2-Exp has 49.

Related comparisons

Go deeper