Model comparison

Command A vs Dolly 2.0-12b

Command A is the stronger model overall, scoring 36.5 to 25.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Command A scores higher in 6 categories and Dolly 2.0-12b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Command A leads 47.6 to 15.2.

Side by side

Command A and Dolly 2.0-12b specifications
Command ADolly 2.0-12b
ProviderCohereDatabricks
Noometry Index36.525.5
Released2025-03-132023-04-11
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Command A leads

Command A: 27.2 (#322), Dolly 2.0-12b: 23.4 (#332)

Coding benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Coding1330776
Aider Polyglot12%—

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Dolly 2.0-12b: —

Agentic & Tool Use benchmarks
BenchmarkCommand ADolly 2.0-12b
Berkeley Function Calling Leaderboard57.1%—

Reasoning Command A leads

Command A: 18.3 (#283), Dolly 2.0-12b: 15.3 (#316)

Reasoning benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Hard Prompts1326804
Kagi LLM Benchmark28.8%—
DTBench61.3%—
LMCA10.3%—
Epoch Capabilities Index—89.67
HellaSwag—70.8%
PIQA—75.4%
WinoGrande—61.8%

Math Command A leads

Command A: 36.2 (#171), Dolly 2.0-12b: 27.3 (#251)

Math benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Math1300871

Knowledge Not comparable

Command A: 37.1 (#159), Dolly 2.0-12b: —

Knowledge benchmarks
BenchmarkCommand ADolly 2.0-12b
Vectara Hallucination Rate9.3%—
LMArena Expert1295—
ARC (AI2) Challenge—39.6%
BoolQ—56.3%
MMLU—26.2%
OpenBookQA—39.2%

Multilingual Command A leads

Command A: 45.3 (#170), Dolly 2.0-12b: 17.4 (#296)

Multilingual benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Non-English1313836
LMArena Chinese1327836
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Russian1314—
LMArena Spanish1347—

Instruction Following Command A leads

Command A: 69.1 (#177), Dolly 2.0-12b: 38.7 (#304)

Instruction Following benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Instruction Following1309814

Long Context Not comparable

Command A: 40.6 (#151), Dolly 2.0-12b: —

Long Context benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Longer Query1334—

Writing & Preference Command A leads

Command A: 47.6 (#208), Dolly 2.0-12b: 15.2 (#311)

Writing & Preference benchmarks
BenchmarkCommand ADolly 2.0-12b
LMArena Text1331851
LMArena Creative Writing1319864
LMArena Multi-Turn1339740
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Dolly 2.0-12b?

Command A is the stronger model overall, scoring 36.5 to 25.5 on the Noometry Index.

Is Command A or Dolly 2.0-12b better for coding?

Command A scores higher on coding benchmarks: 27.2 versus 23.4 in the Noometry coding category.

How many benchmarks do Command A and Dolly 2.0-12b share?

9 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Dolly 2.0-12b has 17.

Related comparisons

Go deeper