Model comparison

Command A vs Llama 3.1-70B

Command A is the stronger model overall, scoring 36.5 to 29.6 on the Noometry Index. Llama 3.1-70B costs 11× less per token, which makes it the better buy when Command A's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Command A scores higher in 7 categories and Llama 3.1-70B in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Command A leads 36.2 to 13.5.
  • Llama 3.1-70B is cheaper at $0.40 / $0.40 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A accepts more context: 256K tokens versus 128K.

Side by side

Command A and Llama 3.1-70B specifications
Command ALlama 3.1-70B
ProviderCohereMeta
Noometry Index36.529.6
Released2025-03-132024-07-23
WeightsOpenOpen
Context window256K128K
Max output8K4K
Input $ / M tokens$2.50$0.40
Output $ / M tokens$10$0.40
Results tracked2435

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-70B leads

Command A: 27.2 (#322), Llama 3.1-70B: 30.3 (#296)

Coding benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Coding13301260
Aider Polyglot12%—
WeirdML—9%
BigCodeBench Instruct—46.1%
BigCodeBench Complete—54.8%

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), Llama 3.1-70B: 25.1 (#112)

Agentic & Tool Use benchmarks
BenchmarkCommand ALlama 3.1-70B
Berkeley Function Calling Leaderboard57.1%—
TheAgentCompany—6.9%
BALROG—27.9%

Reasoning Llama 3.1-70B leads

Command A: 18.3 (#283), Llama 3.1-70B: 21.6 (#220)

Reasoning benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Hard Prompts13261241
DTBench61.3%60%
LMCA10.3%14.8%
Kagi LLM Benchmark28.8%—
Epoch Capabilities Index—125.92

Math Command A leads

Command A: 36.2 (#171), Llama 3.1-70B: 13.5 (#304)

Math benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Math13001252
OTIS Mock AIME 2024-2025—3.6%
Omni-MATH—21%
MATH Level 5—36.7%

Knowledge Command A leads

Command A: 37.1 (#159), Llama 3.1-70B: 24.2 (#269)

Knowledge benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Expert12951209
GPQA Diamond—44.2%
MMLU-Pro—65.3%
Vectara Hallucination Rate9.3%—
GPQA (HELM)—42.6%
MMLU—80.1%

Multilingual Command A leads

Command A: 45.3 (#170), Llama 3.1-70B: 38.8 (#225)

Multilingual benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Non-English13131219
LMArena Chinese13271215
LMArena French13511261
LMArena German13411222
LMArena Japanese12851132
LMArena Korean12851140
LMArena Russian13141234
LMArena Spanish13471253

Instruction Following Command A leads

Command A: 69.1 (#177), Llama 3.1-70B: 65.3 (#223)

Instruction Following benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Instruction Following13091231
IFEval—82.1%

Long Context Command A leads

Command A: 40.6 (#151), Llama 3.1-70B: 37.6 (#214)

Long Context benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Longer Query13341241

Writing & Preference Command A leads

Command A: 47.6 (#208), Llama 3.1-70B: 35.4 (#267)

Writing & Preference benchmarks
BenchmarkCommand ALlama 3.1-70B
LMArena Text13311261
LMArena Creative Writing13191232
EQ-Bench Creative Writing1145784
LMArena Multi-Turn13391256
WildBench—75.8%

Frequently asked questions

Is Command A better than Llama 3.1-70B?

Command A is the stronger model overall, scoring 36.5 to 29.6 on the Noometry Index. Llama 3.1-70B costs 11× less per token, which makes it the better buy when Command A's lead doesn't matter for your workload.

Which is cheaper, Command A or Llama 3.1-70B?

Llama 3.1-70B is cheaper. It lists at $0.40 per million input tokens and $0.40 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or Llama 3.1-70B better for coding?

Llama 3.1-70B scores higher on coding benchmarks: 30.3 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Command A does, with 256K tokens against 128K.

How many benchmarks do Command A and Llama 3.1-70B share?

20 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Llama 3.1-70B has 35.

Related comparisons

Go deeper