Model comparison

Command A vs Gemma 2 9B

Command A is the stronger model overall, scoring 36.5 to 25.9 on the Noometry Index.

Last verified . 18 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Command A scores higher in 7 categories and Gemma 2 9B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command A leads 37.1 to 9.7.

Side by side

Command A and Gemma 2 9B specifications
Command AGemma 2 9B
ProviderCohereGoogle
Noometry Index36.525.9
Released2025-03-132024-06-24
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2435

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 9B leads

Command A: 27.2 (#322), Gemma 2 9B: 29.4 (#304)

Coding benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Coding13301173
Aider Polyglot12%—
BigCodeBench Instruct—34.7%
LiveBench Coding—22.5%
BigCodeBench Complete—40.6%

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Gemma 2 9B: —

Agentic & Tool Use benchmarks
BenchmarkCommand AGemma 2 9B
Berkeley Function Calling Leaderboard57.1%—

Reasoning Command A leads

Command A: 18.3 (#283), Gemma 2 9B: 15.9 (#309)

Reasoning benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Hard Prompts13261171
Kagi LLM Benchmark28.8%—
LiveBench Reasoning—15.2%
DTBench61.3%—
LiveBench Data Analysis—36.4%
LMCA10.3%—
Epoch Capabilities Index—119.83
LiveBench—28.7%
PIQA—83.7%

Math Command A leads

Command A: 36.2 (#171), Gemma 2 9B: 9.9 (#318)

Math benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Math13001183
OTIS Mock AIME 2024-2025—0.6%
LiveBench Math—19.8%
MATH Level 5—21%
GSM8K—84.9%

Knowledge Command A leads

Command A: 37.1 (#159), Gemma 2 9B: 9.7 (#305)

Knowledge benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Expert12951147
GPQA Diamond—27.5%
Vectara Hallucination Rate9.3%—
BoolQ—85.7%
MMLU—72.1%

Multilingual Command A leads

Command A: 45.3 (#170), Gemma 2 9B: 36.6 (#238)

Multilingual benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Non-English13131188
LMArena Chinese13271185
LMArena French13511190
LMArena German13411186
LMArena Japanese12851144
LMArena Korean12851137
LMArena Russian13141200
LMArena Spanish13471200

Instruction Following Command A leads

Command A: 69.1 (#177), Gemma 2 9B: 57.6 (#269)

Instruction Following benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Instruction Following13091178
LiveBench Instruction Following—52.6%

Long Context Command A leads

Command A: 40.6 (#151), Gemma 2 9B: 36.3 (#233)

Long Context benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Longer Query13341197

Writing & Preference Command A leads

Command A: 47.6 (#208), Gemma 2 9B: 32.1 (#281)

Writing & Preference benchmarks
BenchmarkCommand AGemma 2 9B
LMArena Text13311207
LMArena Creative Writing13191206
EQ-Bench Creative Writing1145841
LMArena Multi-Turn13391193
LiveBench Language—25.5%

Frequently asked questions

Is Command A better than Gemma 2 9B?

Command A is the stronger model overall, scoring 36.5 to 25.9 on the Noometry Index.

Is Command A or Gemma 2 9B better for coding?

Gemma 2 9B scores higher on coding benchmarks: 29.4 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Gemma 2 9B share?

18 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Gemma 2 9B has 35.

Related comparisons

Go deeper