Model comparison

Command A vs Llama 3.1 Tulu 3 8b

Command A and Llama 3.1 Tulu 3 8b score almost the same on the Noometry Index (36.5 vs 35.7), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Command A scores higher in 5 categories and Llama 3.1 Tulu 3 8b in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Command A leads 45.3 to 35.4.

Side by side

Command A and Llama 3.1 Tulu 3 8b specifications
Command ALlama 3.1 Tulu 3 8b
ProviderCohereAllen Institute for AI (Ai2)
Noometry Index36.535.7
Released2025-03-13—
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2411

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Tulu 3 8b leads

Command A: 27.2 (#322), Llama 3.1 Tulu 3 8b: 34.4 (#235)

Coding benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Coding13301183
Aider Polyglot12%—

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Llama 3.1 Tulu 3 8b: —

Agentic & Tool Use benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
Berkeley Function Calling Leaderboard57.1%—

Reasoning Llama 3.1 Tulu 3 8b leads

Command A: 18.3 (#283), Llama 3.1 Tulu 3 8b: 22.8 (#188)

Reasoning benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Hard Prompts13261174
Kagi LLM Benchmark28.8%—
DTBench61.3%—
LMCA10.3%—

Math Command A leads

Command A: 36.2 (#171), Llama 3.1 Tulu 3 8b: 33.9 (#198)

Math benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Math13001195

Knowledge Not comparable

Command A: 37.1 (#159), Llama 3.1 Tulu 3 8b: —

Knowledge benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
Vectara Hallucination Rate9.3%—
LMArena Expert1295—

Multilingual Command A leads

Command A: 45.3 (#170), Llama 3.1 Tulu 3 8b: 35.4 (#246)

Multilingual benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Non-English13131169
LMArena Chinese13271176
LMArena Russian13141193
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Spanish1347—

Instruction Following Command A leads

Command A: 69.1 (#177), Llama 3.1 Tulu 3 8b: 61.3 (#246)

Instruction Following benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Instruction Following13091174

Long Context Command A leads

Command A: 40.6 (#151), Llama 3.1 Tulu 3 8b: 35.8 (#239)

Long Context benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Longer Query13341181

Writing & Preference Command A leads

Command A: 47.6 (#208), Llama 3.1 Tulu 3 8b: 39.7 (#245)

Writing & Preference benchmarks
BenchmarkCommand ALlama 3.1 Tulu 3 8b
LMArena Text13311193
LMArena Creative Writing13191182
LMArena Multi-Turn13391154
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Llama 3.1 Tulu 3 8b?

Command A and Llama 3.1 Tulu 3 8b score almost the same on the Noometry Index (36.5 vs 35.7), so choose on price, context window or the category you care about most.

Is Command A or Llama 3.1 Tulu 3 8b better for coding?

Llama 3.1 Tulu 3 8b scores higher on coding benchmarks: 34.4 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Llama 3.1 Tulu 3 8b share?

11 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Llama 3.1 Tulu 3 8b has 11.

Related comparisons

Go deeper