Model comparison

Command A vs Gemma 3 4B

Command A is the stronger model overall, scoring 36.5 to 28.1 on the Noometry Index. Gemma 3 4B costs 88× less per token, which makes it the better buy when Command A's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Gemma 3 4B Google

28.1

Rank #326 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Command A scores higher in 8 categories and Gemma 3 4B in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command A leads 37.1 to 11.8.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 57.1% for Command A and 19.6% for Gemma 3 4B.
  • Gemma 3 4B is cheaper at $0.04 / $0.08 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A accepts more context: 256K tokens versus 131K.

Side by side

Command A and Gemma 3 4B specifications
Command AGemma 3 4B
ProviderCohereGoogle
Noometry Index36.528.1
Released2025-03-132025-03-12
WeightsOpenOpen
Context window256K131K
Max output8K4K
Input $ / M tokens$2.50$0.04
Output $ / M tokens$10$0.08
Results tracked2422

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3 4B leads

Command A: 27.2 (#322), Gemma 3 4B: 35.9 (#215)

Coding benchmarks
BenchmarkCommand AGemma 3 4B
LMArena Coding13301230
Aider Polyglot12%—

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), Gemma 3 4B: 20.9 (#142)

Agentic & Tool Use benchmarks
BenchmarkCommand AGemma 3 4B
Berkeley Function Calling Leaderboard57.1%19.6%

Reasoning Command A leads

Command A: 18.3 (#283), Gemma 3 4B: 13.2 (#335)

Reasoning benchmarks
BenchmarkCommand AGemma 3 4B
Kagi LLM Benchmark28.8%25.2%
LMArena Hard Prompts13261253
DTBench61.3%50.9%
LMCA10.3%2.8%
Chess Puzzles—0%
Epoch Capabilities Index—116.02

Math Command A leads

Command A: 36.2 (#171), Gemma 3 4B: 16.8 (#292)

Math benchmarks
BenchmarkCommand AGemma 3 4B
LMArena Math13001239
OTIS Mock AIME 2024-2025—7.5%

Knowledge Command A leads

Command A: 37.1 (#159), Gemma 3 4B: 11.8 (#299)

Knowledge benchmarks
BenchmarkCommand AGemma 3 4B
Vectara Hallucination Rate9.3%6.4%
LMArena Expert12951223
GPQA Diamond—23.2%

Multilingual Command A leads

Command A: 45.3 (#170), Gemma 3 4B: 42.5 (#194)

Multilingual benchmarks
BenchmarkCommand AGemma 3 4B
LMArena Non-English13131273
LMArena German13411281
LMArena Russian13141294
LMArena Chinese1327—
LMArena French1351—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Spanish1347—

Instruction Following Command A leads

Command A: 69.1 (#177), Gemma 3 4B: 65.2 (#225)

Instruction Following benchmarks
BenchmarkCommand AGemma 3 4B
LMArena Instruction Following13091239

Long Context Command A leads

Command A: 40.6 (#151), Gemma 3 4B: 38.7 (#194)

Long Context benchmarks
BenchmarkCommand AGemma 3 4B
LMArena Longer Query13341273

Writing & Preference Command A leads

Command A: 47.6 (#208), Gemma 3 4B: 42.0 (#239)

Writing & Preference benchmarks
BenchmarkCommand AGemma 3 4B
LMArena Text13311291
LMArena Creative Writing13191271
EQ-Bench Creative Writing11451068
LMArena Multi-Turn13391255

Frequently asked questions

Is Command A better than Gemma 3 4B?

Command A is the stronger model overall, scoring 36.5 to 28.1 on the Noometry Index. Gemma 3 4B costs 88× less per token, which makes it the better buy when Command A's lead doesn't matter for your workload.

Which is cheaper, Command A or Gemma 3 4B?

Gemma 3 4B is cheaper. It lists at $0.04 per million input tokens and $0.08 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or Gemma 3 4B better for coding?

Gemma 3 4B scores higher on coding benchmarks: 35.9 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Command A does, with 256K tokens against 131K.

How many benchmarks do Command A and Gemma 3 4B share?

18 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Gemma 3 4B has 22.

Related comparisons

Go deeper