Model comparison

Command A vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 36.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Command A scores higher in 5 categories and Mercury in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mercury leads 38.7 to 27.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 28.8% for Command A and 21.6% for Mercury.
  • Command A has downloadable open weights; the other is API-only.

Side by side

Command A and Mercury specifications
Command AMercury
ProviderCohereInception
Noometry Index36.537.6
Released2025-03-13—
WeightsOpenProprietary
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked249

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Command A: 27.2 (#322), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkCommand AMercury
LMArena Coding13301322
Aider Polyglot12%—

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkCommand AMercury
Berkeley Function Calling Leaderboard57.1%—

Reasoning Too close to call

Command A: 18.3 (#283), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkCommand AMercury
Kagi LLM Benchmark28.8%21.6%
LMArena Hard Prompts13261285
DTBench61.3%—
LMCA10.3%—

Math Not comparable

Command A: 36.2 (#171), Mercury: —

Math benchmarks
BenchmarkCommand AMercury
LMArena Math1300—

Knowledge Not comparable

Command A: 37.1 (#159), Mercury: —

Knowledge benchmarks
BenchmarkCommand AMercury
Vectara Hallucination Rate9.3%—
LMArena Expert1295—

Multilingual Command A leads

Command A: 45.3 (#170), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkCommand AMercury
LMArena Non-English13131260
LMArena Chinese1327—
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Russian1314—
LMArena Spanish1347—

Instruction Following Command A leads

Command A: 69.1 (#177), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkCommand AMercury
LMArena Instruction Following13091239

Long Context Command A leads

Command A: 40.6 (#151), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkCommand AMercury
LMArena Longer Query13341266

Writing & Preference Command A leads

Command A: 47.6 (#208), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkCommand AMercury
LMArena Text13311282
LMArena Creative Writing13191191
LMArena Multi-Turn13391282
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 36.5 on the Noometry Index.

Is Command A or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Mercury share?

9 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper