Model comparison

Command R vs Granite 3.1 2b Instruct

Granite 3.1 2b Instruct is the stronger model overall, scoring 33.2 to 31.4 on the Noometry Index.

Last verified . 12 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Granite 3.1 2b Instruct IBM

33.2

Rank #247 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Command R scores higher in 5 categories and Granite 3.1 2b Instruct in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Granite 3.1 2b Instruct leads 22.0 to 13.8.

Side by side

Command R and Granite 3.1 2b Instruct specifications
Command RGranite 3.1 2b Instruct
ProviderCohereIBM
Noometry Index31.433.2
Released2024-08-30—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2912

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 3.1 2b Instruct leads

Command R: 29.3 (#306), Granite 3.1 2b Instruct: 33.4 (#257)

Coding benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Coding11691149
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Reasoning Granite 3.1 2b Instruct leads

Command R: 13.8 (#331), Granite 3.1 2b Instruct: 22.0 (#209)

Reasoning benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Hard Prompts11641138
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
LiveBench27.5%—

Math Granite 3.1 2b Instruct leads

Command R: 28.0 (#246), Granite 3.1 2b Instruct: 33.1 (#206)

Math benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Math11551159
LiveBench Math19.4%—

Knowledge Too close to call

Command R: 31.0 (#221), Granite 3.1 2b Instruct: 30.8 (#224)

Knowledge benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Expert11381131
MMLU65.2%—

Multilingual Command R leads

Command R: 35.7 (#245), Granite 3.1 2b Instruct: 29.1 (#269)

Multilingual benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Non-English11741068
LMArena Chinese11821139
LMArena Russian11741063
LMArena French1162—
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Spanish1151—

Instruction Following Too close to call

Command R: 58.1 (#261), Granite 3.1 2b Instruct: 57.7 (#264)

Instruction Following benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Instruction Following11671116
LiveBench Instruction Following55.6%—

Long Context Command R leads

Command R: 36.3 (#231), Granite 3.1 2b Instruct: 35.0 (#244)

Long Context benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Longer Query11981155

Writing & Preference Command R leads

Command R: 38.2 (#254), Granite 3.1 2b Instruct: 34.1 (#274)

Writing & Preference benchmarks
BenchmarkCommand RGranite 3.1 2b Instruct
LMArena Text11871127
LMArena Creative Writing11701116
LMArena Multi-Turn11631099
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Granite 3.1 2b Instruct?

Granite 3.1 2b Instruct is the stronger model overall, scoring 33.2 to 31.4 on the Noometry Index.

Is Command R or Granite 3.1 2b Instruct better for coding?

Granite 3.1 2b Instruct scores higher on coding benchmarks: 33.4 versus 29.3 in the Noometry coding category.

How many benchmarks do Command R and Granite 3.1 2b Instruct share?

12 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Granite 3.1 2b Instruct has 12.

Related comparisons

Go deeper