Model comparison

Command A vs Llama 13b

Command A is the stronger model overall, scoring 36.5 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Command A scores higher in 6 categories and Llama 13b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Command A leads 47.6 to 13.8.

Side by side

Command A and Llama 13b specifications
Command ALlama 13b
ProviderCohereMeta
Noometry Index36.524.4
Released2025-03-132023-02-24
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2421

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Command A leads

Command A: 27.2 (#322), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkCommand ALlama 13b
LMArena Coding1330683
Aider Polyglot12%—

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Llama 13b: —

Agentic & Tool Use benchmarks
BenchmarkCommand ALlama 13b
Berkeley Function Calling Leaderboard57.1%—

Reasoning Command A leads

Command A: 18.3 (#283), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkCommand ALlama 13b
LMArena Hard Prompts1326728
Kagi LLM Benchmark28.8%—
DTBench61.3%—
LMCA10.3%—
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Command A leads

Command A: 36.2 (#171), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkCommand ALlama 13b
LMArena Math1300838
GSM8K—20.6%

Knowledge Not comparable

Command A: 37.1 (#159), Llama 13b: —

Knowledge benchmarks
BenchmarkCommand ALlama 13b
Vectara Hallucination Rate9.3%—
LMArena Expert1295—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Command A: —, Llama 13b: —

Multimodal benchmarks
BenchmarkCommand ALlama 13b
ScienceQA—43.3%

Multilingual Command A leads

Command A: 45.3 (#170), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkCommand ALlama 13b
LMArena Non-English1313819
LMArena Chinese1327—
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Russian1314—
LMArena Spanish1347—

Instruction Following Command A leads

Command A: 69.1 (#177), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkCommand ALlama 13b
LMArena Instruction Following1309781

Long Context Not comparable

Command A: 40.6 (#151), Llama 13b: —

Long Context benchmarks
BenchmarkCommand ALlama 13b
LMArena Longer Query1334—

Writing & Preference Command A leads

Command A: 47.6 (#208), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkCommand ALlama 13b
LMArena Text1331834
LMArena Creative Writing1319794
LMArena Multi-Turn1339753
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Llama 13b?

Command A is the stronger model overall, scoring 36.5 to 24.4 on the Noometry Index.

Is Command A or Llama 13b better for coding?

Command A scores higher on coding benchmarks: 27.2 versus 21.4 in the Noometry coding category.

How many benchmarks do Command A and Llama 13b share?

8 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper