Model comparison

Command A vs Qwen1.5-32B

Command A is the stronger model overall, scoring 36.5 to 30.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Qwen1.5-32B Alibaba (Qwen)

30.5

Rank #293 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Command A scores higher in 6 categories and Qwen1.5-32B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command A leads 37.1 to 13.5.

Side by side

Command A and Qwen1.5-32B specifications
Command AQwen1.5-32B
ProviderCohereAlibaba (Qwen)
Noometry Index36.530.5
Released2025-03-132024-02-04
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2421

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-32B leads

Command A: 27.2 (#322), Qwen1.5-32B: 31.7 (#282)

Coding benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Coding13301155
Aider Polyglot12%—
BigCodeBench Instruct—32.3%
BigCodeBench Complete—42%

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Qwen1.5-32B: —

Agentic & Tool Use benchmarks
BenchmarkCommand AQwen1.5-32B
Berkeley Function Calling Leaderboard57.1%—

Reasoning Qwen1.5-32B leads

Command A: 18.3 (#283), Qwen1.5-32B: 21.8 (#212)

Reasoning benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Hard Prompts13261130
Kagi LLM Benchmark28.8%—
DTBench61.3%—
LMCA10.3%—

Math Command A leads

Command A: 36.2 (#171), Qwen1.5-32B: 33.0 (#207)

Math benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Math13001155

Knowledge Command A leads

Command A: 37.1 (#159), Qwen1.5-32B: 13.5 (#296)

Knowledge benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Expert12951126
GPQA Diamond—30.7%
Vectara Hallucination Rate9.3%—
MMLU—74.4%

Multilingual Command A leads

Command A: 45.3 (#170), Qwen1.5-32B: 31.4 (#259)

Multilingual benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Non-English13131106
LMArena Chinese13271177
LMArena French13511101
LMArena German13411058
LMArena Japanese12851027
LMArena Korean12851008
LMArena Russian13141073
LMArena Spanish13471089

Instruction Following Command A leads

Command A: 69.1 (#177), Qwen1.5-32B: 57.7 (#265)

Instruction Following benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Instruction Following13091116

Long Context Command A leads

Command A: 40.6 (#151), Qwen1.5-32B: 34.7 (#246)

Long Context benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Longer Query13341146

Writing & Preference Command A leads

Command A: 47.6 (#208), Qwen1.5-32B: 34.2 (#271)

Writing & Preference benchmarks
BenchmarkCommand AQwen1.5-32B
LMArena Text13311137
LMArena Creative Writing13191083
LMArena Multi-Turn13391140
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Qwen1.5-32B?

Command A is the stronger model overall, scoring 36.5 to 30.5 on the Noometry Index.

Is Command A or Qwen1.5-32B better for coding?

Qwen1.5-32B scores higher on coding benchmarks: 31.7 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Qwen1.5-32B share?

17 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Qwen1.5-32B has 21.

Related comparisons

Go deeper