Model comparison

Command R vs Llama 3.2 1B

Command R is the stronger model overall, scoring 31.4 to 20.1 on the Noometry Index. Llama 3.2 1B costs 3.7× less per token, which makes it the better buy when Command R's lead doesn't matter for your workload.

Last verified . 15 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Command R scores higher in 7 categories and Llama 3.2 1B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command R leads 31.0 to 7.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 45.2% for Command R and 11.3% for Llama 3.2 1B.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.15 / $0.60 for Command R.
  • Command R accepts more context: 128K tokens versus 60K.

Side by side

Command R and Llama 3.2 1B specifications
Command RLlama 3.2 1B
ProviderCohereMeta
Noometry Index31.420.1
Released2024-08-302024-09-24
WeightsOpenOpen
Context window128K60K
Max output4K54K
Input $ / M tokens$0.15$0.027
Output $ / M tokens$0.60$0.20
Results tracked2922

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Command R leads

Command R: 29.3 (#306), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkCommand RLlama 3.2 1B
BigCodeBench Instruct37.1%8.2%
LMArena Coding11691070
BigCodeBench Complete45.2%11.3%
LiveBench Coding17.9%—

Agentic & Tool Use Not comparable

Command R: —, Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkCommand RLlama 3.2 1B
Berkeley Function Calling Leaderboard—10.8%
BALROG—6.6%

Reasoning Llama 3.2 1B leads

Command R: 13.8 (#331), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Hard Prompts11641044
Chess Puzzles—0%
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
Epoch Capabilities Index—101.99
LiveBench27.5%—

Math Command R leads

Command R: 28.0 (#246), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Math11551086
OTIS Mock AIME 2024-2025—0.6%
LiveBench Math19.4%—

Knowledge Command R leads

Command R: 31.0 (#221), Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Expert11381007
GPQA Diamond—23.9%
MMLU65.2%—

Multilingual Command R leads

Command R: 35.7 (#245), Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Non-English1174973
LMArena Chinese1182959
LMArena German11761014
LMArena Russian1174941
LMArena French1162—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Spanish1151—

Instruction Following Command R leads

Command R: 58.1 (#261), Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Instruction Following11671031
LiveBench Instruction Following55.6%—

Long Context Command R leads

Command R: 36.3 (#231), Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Longer Query11981050

Writing & Preference Command R leads

Command R: 38.2 (#254), Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkCommand RLlama 3.2 1B
LMArena Text11871055
LMArena Creative Writing11701033
LMArena Multi-Turn11631030
EQ-Bench Creative Writing—200
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Llama 3.2 1B?

Command R is the stronger model overall, scoring 31.4 to 20.1 on the Noometry Index. Llama 3.2 1B costs 3.7× less per token, which makes it the better buy when Command R's lead doesn't matter for your workload.

Which is cheaper, Command R or Llama 3.2 1B?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Command R lists at $0.15 and $0.60.

Is Command R or Llama 3.2 1B better for coding?

Command R scores higher on coding benchmarks: 29.3 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Command R does, with 128K tokens against 60K.

How many benchmarks do Command R and Llama 3.2 1B share?

15 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper