Model comparison

Command R vs Qwen-14B

Command R and Qwen-14B score almost the same on the Noometry Index (31.4 vs 31.4), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Command R scores higher in 4 categories and Qwen-14B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Command R leads 38.2 to 27.6.

Side by side

Command R and Qwen-14B specifications
Command RQwen-14B
ProviderCohereAlibaba (Qwen)
Noometry Index31.431.4
Released2024-08-302023-09-24
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen-14B leads

Command R: 29.3 (#306), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkCommand RQwen-14B
LMArena Coding11691071
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Reasoning Qwen-14B leads

Command R: 13.8 (#331), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkCommand RQwen-14B
LMArena Hard Prompts11641027
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
BIG-Bench Hard—55%
Epoch Capabilities Index—113.03
LAMBADA—71.1%
LiveBench27.5%—
PIQA—79.9%

Math Qwen-14B leads

Command R: 28.0 (#246), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkCommand RQwen-14B
LMArena Math11551068
LiveBench Math19.4%—
GSM8K—61.3%

Knowledge Not comparable

Command R: 31.0 (#221), Qwen-14B: —

Knowledge benchmarks
BenchmarkCommand RQwen-14B
MMLU65.2%66.3%
LMArena Expert1138—
ARC (AI2) Challenge—84.4%
BoolQ—86.2%

Multilingual Command R leads

Command R: 35.7 (#245), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkCommand RQwen-14B
LMArena Non-English11741041
LMArena Chinese11821077
LMArena French1162—
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Russian1174—
LMArena Spanish1151—

Instruction Following Command R leads

Command R: 58.1 (#261), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkCommand RQwen-14B
LMArena Instruction Following11671031
LiveBench Instruction Following55.6%—

Long Context Command R leads

Command R: 36.3 (#231), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkCommand RQwen-14B
LMArena Longer Query11981028

Writing & Preference Command R leads

Command R: 38.2 (#254), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkCommand RQwen-14B
LMArena Text11871051
LMArena Creative Writing11701028
LMArena Multi-Turn11631022
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Qwen-14B?

Command R and Qwen-14B score almost the same on the Noometry Index (31.4 vs 31.4), so choose on price, context window or the category you care about most.

Is Command R or Qwen-14B better for coding?

Qwen-14B scores higher on coding benchmarks: 31.2 versus 29.3 in the Noometry coding category.

How many benchmarks do Command R and Qwen-14B share?

11 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper