Model comparison

Command R vs Dolly 2.0-12b

Command R is the stronger model overall, scoring 31.4 to 25.5 on the Noometry Index.

Last verified . 10 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Command R scores higher in 5 categories and Dolly 2.0-12b in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Command R leads 38.2 to 15.2.

Side by side

Command R and Dolly 2.0-12b specifications
Command RDolly 2.0-12b
ProviderCohereDatabricks
Noometry Index31.425.5
Released2024-08-302023-04-11
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Command R leads

Command R: 29.3 (#306), Dolly 2.0-12b: 23.4 (#332)

Coding benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Coding1169776
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Reasoning Dolly 2.0-12b leads

Command R: 13.8 (#331), Dolly 2.0-12b: 15.3 (#316)

Reasoning benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Hard Prompts1164804
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
Epoch Capabilities Index—89.67
HellaSwag—70.8%
LiveBench27.5%—
PIQA—75.4%
WinoGrande—61.8%

Math Too close to call

Command R: 28.0 (#246), Dolly 2.0-12b: 27.3 (#251)

Math benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Math1155871
LiveBench Math19.4%—

Knowledge Not comparable

Command R: 31.0 (#221), Dolly 2.0-12b: —

Knowledge benchmarks
BenchmarkCommand RDolly 2.0-12b
MMLU65.2%26.2%
LMArena Expert1138—
ARC (AI2) Challenge—39.6%
BoolQ—56.3%
OpenBookQA—39.2%

Multilingual Command R leads

Command R: 35.7 (#245), Dolly 2.0-12b: 17.4 (#296)

Multilingual benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Non-English1174836
LMArena Chinese1182836
LMArena French1162—
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Russian1174—
LMArena Spanish1151—

Instruction Following Command R leads

Command R: 58.1 (#261), Dolly 2.0-12b: 38.7 (#304)

Instruction Following benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Instruction Following1167814
LiveBench Instruction Following55.6%—

Long Context Not comparable

Command R: 36.3 (#231), Dolly 2.0-12b: —

Long Context benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Longer Query1198—

Writing & Preference Command R leads

Command R: 38.2 (#254), Dolly 2.0-12b: 15.2 (#311)

Writing & Preference benchmarks
BenchmarkCommand RDolly 2.0-12b
LMArena Text1187851
LMArena Creative Writing1170864
LMArena Multi-Turn1163740
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Dolly 2.0-12b?

Command R is the stronger model overall, scoring 31.4 to 25.5 on the Noometry Index.

Is Command R or Dolly 2.0-12b better for coding?

Command R scores higher on coding benchmarks: 29.3 versus 23.4 in the Noometry coding category.

How many benchmarks do Command R and Dolly 2.0-12b share?

10 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Dolly 2.0-12b has 17.

Related comparisons

Go deeper