Model comparison

Command A vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 36.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Command A scores higher in 1 category and Grok 4.1 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 leads 62.4 to 47.6.
  • Command A has downloadable open weights; the other is API-only.

Side by side

Command A and Grok 4.1 specifications
Command AGrok 4.1
ProviderCoherexAI
Noometry Index36.541.5
Released2025-03-132025-11-17
WeightsOpenProprietary
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 leads

Command A: 27.2 (#322), Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkCommand AGrok 4.1
LMArena Coding13301445
Aider Polyglot12%—
LMArena WebDev—1214

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkCommand AGrok 4.1
Berkeley Function Calling Leaderboard57.1%—
Cybench—39%

Reasoning Grok 4.1 leads

Command A: 18.3 (#283), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkCommand AGrok 4.1
LMArena Hard Prompts13261435
Kagi LLM Benchmark28.8%—
DTBench61.3%—
LMCA10.3%—

Math Grok 4.1 leads

Command A: 36.2 (#171), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkCommand AGrok 4.1
LMArena Math13001422

Knowledge Grok 4.1 leads

Command A: 37.1 (#159), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkCommand AGrok 4.1
LMArena Expert12951417
Vectara Hallucination Rate9.3%—

Multilingual Grok 4.1 leads

Command A: 45.3 (#170), Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkCommand AGrok 4.1
LMArena Non-English13131425
LMArena Chinese13271465
LMArena French13511448
LMArena German13411446
LMArena Japanese12851397
LMArena Korean12851407
LMArena Russian13141434
LMArena Spanish13471438

Instruction Following Grok 4.1 leads

Command A: 69.1 (#177), Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkCommand AGrok 4.1
LMArena Instruction Following13091400

Long Context Grok 4.1 leads

Command A: 40.6 (#151), Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkCommand AGrok 4.1
LMArena Longer Query13341416

Writing & Preference Grok 4.1 leads

Command A: 47.6 (#208), Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkCommand AGrok 4.1
LMArena Text13311437
LMArena Creative Writing13191411
LMArena Multi-Turn13391437
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 36.5 on the Noometry Index.

Is Command A or Grok 4.1 better for coding?

Grok 4.1 scores higher on coding benchmarks: 33.7 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Grok 4.1 share?

17 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper