Model comparison

Command R vs Llama 3.1-70B

Command R is the stronger model overall, scoring 31.4 to 29.6 on the Noometry Index.

Last verified . 22 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Command R scores higher in 3 categories and Llama 3.1-70B in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Command R leads 28.0 to 13.5.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 60% for Llama 3.1-70B.
  • Command R is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.40 / $0.40 for Llama 3.1-70B.

Side by side

Command R and Llama 3.1-70B specifications
Command RLlama 3.1-70B
ProviderCohereMeta
Noometry Index31.429.6
Released2024-08-302024-07-23
WeightsOpenOpen
Context window128K128K
Max output4K4K
Input $ / M tokens$0.15$0.40
Output $ / M tokens$0.60$0.40
Results tracked2935

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Command R: 29.3 (#306), Llama 3.1-70B: 30.3 (#296)

Coding benchmarks
BenchmarkCommand RLlama 3.1-70B
BigCodeBench Instruct37.1%46.1%
LMArena Coding11691260
BigCodeBench Complete45.2%54.8%
WeirdML—9%
LiveBench Coding17.9%—

Agentic & Tool Use Not comparable

Command R: —, Llama 3.1-70B: 25.1 (#112)

Agentic & Tool Use benchmarks
BenchmarkCommand RLlama 3.1-70B
TheAgentCompany—6.9%
BALROG—27.9%

Reasoning Llama 3.1-70B leads

Command R: 13.8 (#331), Llama 3.1-70B: 21.6 (#220)

Reasoning benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Hard Prompts11641241
DTBench46.4%60%
LMCA9.2%14.8%
LiveBench Reasoning21.9%—
LiveBench Data Analysis33.3%—
Epoch Capabilities Index—125.92
LiveBench27.5%—

Math Command R leads

Command R: 28.0 (#246), Llama 3.1-70B: 13.5 (#304)

Math benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Math11551252
OTIS Mock AIME 2024-2025—3.6%
Omni-MATH—21%
LiveBench Math19.4%—
MATH Level 5—36.7%

Knowledge Command R leads

Command R: 31.0 (#221), Llama 3.1-70B: 24.2 (#269)

Knowledge benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Expert11381209
MMLU65.2%80.1%
GPQA Diamond—44.2%
MMLU-Pro—65.3%
GPQA (HELM)—42.6%

Multilingual Llama 3.1-70B leads

Command R: 35.7 (#245), Llama 3.1-70B: 38.8 (#225)

Multilingual benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Non-English11741219
LMArena Chinese11821215
LMArena French11621261
LMArena German11761222
LMArena Japanese11431132
LMArena Korean11631140
LMArena Russian11741234
LMArena Spanish11511253

Instruction Following Llama 3.1-70B leads

Command R: 58.1 (#261), Llama 3.1-70B: 65.3 (#223)

Instruction Following benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Instruction Following11671231
LiveBench Instruction Following55.6%—
IFEval—82.1%

Long Context Llama 3.1-70B leads

Command R: 36.3 (#231), Llama 3.1-70B: 37.6 (#214)

Long Context benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Longer Query11981241

Writing & Preference Command R leads

Command R: 38.2 (#254), Llama 3.1-70B: 35.4 (#267)

Writing & Preference benchmarks
BenchmarkCommand RLlama 3.1-70B
LMArena Text11871261
LMArena Creative Writing11701232
LMArena Multi-Turn11631256
EQ-Bench Creative Writing—784
WildBench—75.8%
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Llama 3.1-70B?

Command R is the stronger model overall, scoring 31.4 to 29.6 on the Noometry Index.

Which is cheaper, Command R or Llama 3.1-70B?

Command R is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Llama 3.1-70B lists at $0.40 and $0.40.

Is Command R or Llama 3.1-70B better for coding?

They score almost the same on coding (29.3 vs 30.3); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Command R and Llama 3.1-70B share?

22 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Llama 3.1-70B has 35.

Related comparisons

Go deeper