Model comparison

Command R vs Phi 3 Small 8k Instruct

Command R is the stronger model overall, scoring 31.4 to 29.3 on the Noometry Index.

Last verified . 25 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Command R scores higher in 7 categories and Phi 3 Small 8k Instruct in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Command R leads 35.7 to 28.5.
  • The biggest single-benchmark swing is LiveBench Instruction Following: 55.6% for Command R and 47.2% for Phi 3 Small 8k Instruct.

Side by side

Command R and Phi 3 Small 8k Instruct specifications
Command RPhi 3 Small 8k Instruct
ProviderCohereMicrosoft
Noometry Index31.429.3
Released2024-08-302024-04-23
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked2932

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Command R leads

Command R: 29.3 (#306), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LiveBench Coding17.9%20.3%
LMArena Coding11691101
BigCodeBench Instruct37.1%—
BigCodeBench Complete45.2%—

Reasoning Too close to call

Command R: 13.8 (#331), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LiveBench Reasoning21.9%15.9%
LMArena Hard Prompts11641100
LiveBench Data Analysis33.3%30.3%
LiveBench27.5%24%
DTBench46.4%—
LMCA9.2%—
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
HellaSwag—77%
WinoGrande—81.5%

Math Too close to call

Command R: 28.0 (#246), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LiveBench Math19.4%17.6%
LMArena Math11551151

Knowledge Command R leads

Command R: 31.0 (#221), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LMArena Expert11381067
MMLU65.2%75.7%
ARC (AI2) Challenge—90.7%
OpenBookQA—88%
TriviaQA—58.1%

Multilingual Command R leads

Command R: 35.7 (#245), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LMArena Non-English11741058
LMArena Chinese11821061
LMArena French11621135
LMArena German11761080
LMArena Japanese1143966
LMArena Korean1163894
LMArena Russian11741111
LMArena Spanish11511111

Instruction Following Command R leads

Command R: 58.1 (#261), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LiveBench Instruction Following55.6%47.2%
LMArena Instruction Following11671087

Long Context Command R leads

Command R: 36.3 (#231), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LMArena Longer Query11981088

Writing & Preference Command R leads

Command R: 38.2 (#254), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkCommand RPhi 3 Small 8k Instruct
LMArena Text11871110
LMArena Creative Writing11701083
LMArena Multi-Turn11631068
LiveBench Language16.7%12.9%

Frequently asked questions

Is Command R better than Phi 3 Small 8k Instruct?

Command R is the stronger model overall, scoring 31.4 to 29.3 on the Noometry Index.

Is Command R or Phi 3 Small 8k Instruct better for coding?

Command R scores higher on coding benchmarks: 29.3 versus 27.9 in the Noometry coding category.

How many benchmarks do Command R and Phi 3 Small 8k Instruct share?

25 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper