Model comparison

Command A vs DeepSeek-R1-Distill-Qwen-32B

Command A and DeepSeek-R1-Distill-Qwen-32B score almost the same on the Noometry Index (36.5 vs 35.5), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

DeepSeek-R1-Distill-Qwen-32B DeepSeek

35.5

Rank #226 Confirmed

Summary

  • The widest gap is in coding, where DeepSeek-R1-Distill-Qwen-32B leads 36.1 to 27.2.

Side by side

Command A and DeepSeek-R1-Distill-Qwen-32B specifications
Command ADeepSeek-R1-Distill-Qwen-32B
ProviderCohereDeepSeek
Noometry Index36.535.5
Released2025-03-132025-01-20
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2414

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1-Distill-Qwen-32B leads

Command A: 27.2 (#322), DeepSeek-R1-Distill-Qwen-32B: 36.1 (#212)

Coding benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
Aider Polyglot12%—
BigCodeBench Instruct—43.9%
LiveBench Coding—33.7%
LMArena Coding1330—
BigCodeBench Complete—54.9%

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), DeepSeek-R1-Distill-Qwen-32B: 28.1 (#94)

Agentic & Tool Use benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
Berkeley Function Calling Leaderboard57.1%—
BALROG—19.5%

Reasoning Too close to call

Command A: 18.3 (#283), DeepSeek-R1-Distill-Qwen-32B: 18.2 (#284)

Reasoning benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
Kagi LLM Benchmark28.8%—
Chess Puzzles—1%
LiveBench Reasoning—52.3%
LMArena Hard Prompts1326—
DTBench61.3%—
LiveBench Data Analysis—45.4%
LMCA10.3%—
Epoch Capabilities Index—137.44
LiveBench—45.5%

Math Command A leads

Command A: 36.2 (#171), DeepSeek-R1-Distill-Qwen-32B: 34.5 (#194)

Math benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
OTIS Mock AIME 2024-2025—55.6%
LiveBench Math—59.4%
LMArena Math1300—

Knowledge Command A leads

Command A: 37.1 (#159), DeepSeek-R1-Distill-Qwen-32B: 35.7 (#182)

Knowledge benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
GPQA Diamond—64.1%
Vectara Hallucination Rate9.3%—
LMArena Expert1295—

Multilingual Not comparable

Command A: 45.3 (#170), DeepSeek-R1-Distill-Qwen-32B: —

Multilingual benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
LMArena Non-English1313—
LMArena Chinese1327—
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Russian1314—
LMArena Spanish1347—

Instruction Following Command A leads

Command A: 69.1 (#177), DeepSeek-R1-Distill-Qwen-32B: 61.6 (#243)

Instruction Following benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
LiveBench Instruction Following—55.7%
LMArena Instruction Following1309—

Long Context Not comparable

Command A: 40.6 (#151), DeepSeek-R1-Distill-Qwen-32B: —

Long Context benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
LMArena Longer Query1334—

Writing & Preference DeepSeek-R1-Distill-Qwen-32B leads

Command A: 47.6 (#208), DeepSeek-R1-Distill-Qwen-32B: 49.6 (#188)

Writing & Preference benchmarks
BenchmarkCommand ADeepSeek-R1-Distill-Qwen-32B
LMArena Text1331—
LMArena Creative Writing1319—
EQ-Bench Creative Writing1145—
LMArena Multi-Turn1339—
LiveBench Language—26.8%

Frequently asked questions

Is Command A better than DeepSeek-R1-Distill-Qwen-32B?

Command A and DeepSeek-R1-Distill-Qwen-32B score almost the same on the Noometry Index (36.5 vs 35.5), so choose on price, context window or the category you care about most.

Is Command A or DeepSeek-R1-Distill-Qwen-32B better for coding?

DeepSeek-R1-Distill-Qwen-32B scores higher on coding benchmarks: 36.1 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and DeepSeek-R1-Distill-Qwen-32B share?

0 benchmarks have published results for both models. Command A has 24 scored results on Noometry and DeepSeek-R1-Distill-Qwen-32B has 14.

Related comparisons

Go deeper