Model comparison

Command R vs Qwen1.5-32B

Command R and Qwen1.5-32B score almost the same on the Noometry Index (31.4 vs 30.5), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Qwen1.5-32B Alibaba (Qwen)

30.5

Rank #293 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Command R scores higher in 5 categories and Qwen1.5-32B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command R leads 31.0 to 13.5.

Side by side

Command R and Qwen1.5-32B specifications
Command RQwen1.5-32B
ProviderCohereAlibaba (Qwen)
Noometry Index31.430.5
Released2024-08-302024-02-04
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2921

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-32B leads

Command R: 29.3 (#306), Qwen1.5-32B: 31.7 (#282)

Coding benchmarks
BenchmarkCommand RQwen1.5-32B
BigCodeBench Instruct37.1%32.3%
LMArena Coding11691155
BigCodeBench Complete45.2%42%
LiveBench Coding17.9%—

Reasoning Qwen1.5-32B leads

Command R: 13.8 (#331), Qwen1.5-32B: 21.8 (#212)

Reasoning benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Hard Prompts11641130
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
LiveBench27.5%—

Math Qwen1.5-32B leads

Command R: 28.0 (#246), Qwen1.5-32B: 33.0 (#207)

Math benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Math11551155
LiveBench Math19.4%—

Knowledge Command R leads

Command R: 31.0 (#221), Qwen1.5-32B: 13.5 (#296)

Knowledge benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Expert11381126
MMLU65.2%74.4%
GPQA Diamond—30.7%

Multilingual Command R leads

Command R: 35.7 (#245), Qwen1.5-32B: 31.4 (#259)

Multilingual benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Non-English11741106
LMArena Chinese11821177
LMArena French11621101
LMArena German11761058
LMArena Japanese11431027
LMArena Korean11631008
LMArena Russian11741073
LMArena Spanish11511089

Instruction Following Too close to call

Command R: 58.1 (#261), Qwen1.5-32B: 57.7 (#265)

Instruction Following benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Instruction Following11671116
LiveBench Instruction Following55.6%—

Long Context Command R leads

Command R: 36.3 (#231), Qwen1.5-32B: 34.7 (#246)

Long Context benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Longer Query11981146

Writing & Preference Command R leads

Command R: 38.2 (#254), Qwen1.5-32B: 34.2 (#271)

Writing & Preference benchmarks
BenchmarkCommand RQwen1.5-32B
LMArena Text11871137
LMArena Creative Writing11701083
LMArena Multi-Turn11631140
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Qwen1.5-32B?

Command R and Qwen1.5-32B score almost the same on the Noometry Index (31.4 vs 30.5), so choose on price, context window or the category you care about most.

Is Command R or Qwen1.5-32B better for coding?

Qwen1.5-32B scores higher on coding benchmarks: 31.7 versus 29.3 in the Noometry coding category.

How many benchmarks do Command R and Qwen1.5-32B share?

20 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Qwen1.5-32B has 21.

Related comparisons

Go deeper