Model comparison

Command A vs Grok Build 0.1

Command A and Grok Build 0.1 score almost the same on the Noometry Index (36.5 vs 36.4), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 27.2.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A has downloadable open weights; the other is API-only.

Side by side

Command A and Grok Build 0.1 specifications
Command AGrok Build 0.1
ProviderCoherexAI
Noometry Index36.536.4
Released2025-03-132026-04-16
WeightsOpenProprietary
Context window256K256K
Max output8K256K
Input $ / M tokens$2.50$1
Output $ / M tokens$10$2
Results tracked243

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Command A: 27.2 (#322), Grok Build 0.1: 43.1 (#91)

Coding benchmarks
BenchmarkCommand AGrok Build 0.1
Aider Polyglot12%—
SciCode—50.2%
LMArena Coding1330—

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), Grok Build 0.1: 22.7 (#129)

Agentic & Tool Use benchmarks
BenchmarkCommand AGrok Build 0.1
Berkeley Function Calling Leaderboard57.1%—
GBAEval—2.4%

Reasoning Grok Build 0.1 leads

Command A: 18.3 (#283), Grok Build 0.1: 32.2 (#77)

Reasoning benchmarks
BenchmarkCommand AGrok Build 0.1
Kagi LLM Benchmark28.8%—
CritPt—9.1%
LMArena Hard Prompts1326—
DTBench61.3%—
LMCA10.3%—

Math Not comparable

Command A: 36.2 (#171), Grok Build 0.1: —

Math benchmarks
BenchmarkCommand AGrok Build 0.1
LMArena Math1300—

Knowledge Not comparable

Command A: 37.1 (#159), Grok Build 0.1: —

Knowledge benchmarks
BenchmarkCommand AGrok Build 0.1
Vectara Hallucination Rate9.3%—
LMArena Expert1295—

Multilingual Not comparable

Command A: 45.3 (#170), Grok Build 0.1: —

Multilingual benchmarks
BenchmarkCommand AGrok Build 0.1
LMArena Non-English1313—
LMArena Chinese1327—
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Russian1314—
LMArena Spanish1347—

Instruction Following Not comparable

Command A: 69.1 (#177), Grok Build 0.1: —

Instruction Following benchmarks
BenchmarkCommand AGrok Build 0.1
LMArena Instruction Following1309—

Long Context Not comparable

Command A: 40.6 (#151), Grok Build 0.1: —

Long Context benchmarks
BenchmarkCommand AGrok Build 0.1
LMArena Longer Query1334—

Writing & Preference Not comparable

Command A: 47.6 (#208), Grok Build 0.1: —

Writing & Preference benchmarks
BenchmarkCommand AGrok Build 0.1
LMArena Text1331—
LMArena Creative Writing1319—
EQ-Bench Creative Writing1145—
LMArena Multi-Turn1339—

Frequently asked questions

Is Command A better than Grok Build 0.1?

Command A and Grok Build 0.1 score almost the same on the Noometry Index (36.5 vs 36.4), so choose on price, context window or the category you care about most.

Which is cheaper, Command A or Grok Build 0.1?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or Grok Build 0.1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Both accept 256K tokens.

How many benchmarks do Command A and Grok Build 0.1 share?

0 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Grok Build 0.1 has 3.

Related comparisons

Go deeper