Cohere, open weights
Command A
Command A by Cohere ranks 215th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.5. Its strongest category is agentic & tool use, where it ranks 40th. API pricing starts at $2.50 per million input tokens and $10 per million output tokens, with a 256K-token context window.
Last verified
Specifications
- Noometry rank
- #215 of 354
- Index score
- 36.5
- Evidence
- Confirmed 24 results
- Provider
Cohere
- Released
- March 13, 2025
- Weights
- Open weights
- Reasoning
- No
- Context window
- 256K
- Max output
- 8K
- Input price
- $2.50 / M
- Output price
- $10 / M
- Blended price
- $4.38 / M
- Output speed
- 28 tokens/s Kagi
- Value
- #192 of 219
- Knowledge cutoff
- June 2024
- Input
- text
- Hugging Face
- CohereForAI/c4ai-command-a-03-2025
Category scores
Each category score combines every public result we have in that category.
- Coding 27.2
- Agentic & Tool Use 35.9
- Reasoning 18.3
- Math 36.2
- Knowledge 37.1
- Multilingual 45.3
- Instruction Following 69.1
- Long Context 40.6
- Writing & Preference 47.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 27.2 | #322 | 2 |
| Agentic & Tool Use | 35.9 | #40 | 1 |
| Reasoning | 18.3 | #283 | 4 |
| Math | 36.2 | #171 | 1 |
| Knowledge | 37.1 | #159 | 2 |
| Multilingual | 45.3 | #170 | 1 |
| Instruction Following | 69.1 | #177 | 1 |
| Long Context | 40.6 | #151 | 1 |
| Writing & Preference | 47.6 | #208 | 4 |
Strengths and weaknesses
Categories where Command A places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 35.9 | +5.6 | #40 of 154, top 26% |
| Knowledge | 37.1 | −0.2 | #159 of 314, top 51% |
| Long Context | 40.6 | −0.3 | #151 of 296, top 52% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 27.2 | −11.5 | #322 of 340, top 95% |
| Reasoning | 18.3 | −5.3 | #283 of 350, top 81% |
| Writing & Preference | 47.6 | −6.2 | #208 of 312, top 67% |
Closest competitors
The models ranked just above and below Command A. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | #211 | 37.0 | $0.18 | 172 | Compare |
| o3-mini | #212 | 36.7 | $1.93 | — | Compare |
| Llama 3.1 Nemotron Ultra 253b v1 | #213 | 36.7 | — | — | Compare |
| Granite 4.0 H Small | #214 | 36.5 | — | — | Compare |
| Grok Build 0.1 | #216 | 36.4 | $1.25 | — | Compare |
| gpt-oss-120b | #217 | 36.3 | $0.0703 | 55 | Compare |
| Mistral Medium | #218 | 36.3 | $3 | 68 | Compare |
| GPT-4.1 | #219 | 35.9 | $3.50 | 116 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 12% | #40 of 44, top 91% | Epoch AI | ||
| LMArena Coding | 1330 | #179 of 294, top 61% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 46.5% | fc | Berkeley Function Calling Leaderboard | ||
| Berkeley Function Calling Leaderboard | 57.1% | #10 of 49, top 21% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 28.8% | #93 of 99, top 94% | Kagi LLM Benchmark | ||
| LMArena Hard Prompts | 1326 | #176 of 297, top 60% | LMArena | 2026-10-08 | |
| DTBench | 61.3% | #114 of 151, top 76% | Epoch AI | ||
| LMCA | 10.3% | #111 of 125, top 89% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1300 | #183 of 285, top 65% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Vectara Hallucination Rate (lower is better) | 9.3% | #43 of 96, top 45% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1295 | #178 of 273, top 66% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1313 | #170 of 297, top 58% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1327 | #179 of 285, top 63% | LMArena | 2026-10-08 | |
| LMArena French | 1351 | #144 of 223, top 65% | LMArena | 2026-10-08 | |
| LMArena German | 1341 | #135 of 231, top 59% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1285 | #130 of 211, top 62% | LMArena | 2026-10-08 | |
| LMArena Korean | 1285 | #135 of 213, top 64% | LMArena | 2026-10-08 | |
| LMArena Russian | 1314 | #169 of 283, top 60% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1347 | #146 of 226, top 65% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1309 | #170 of 298, top 58% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1334 | #160 of 291, top 55% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1331 | #172 of 297, top 58% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1319 | #153 of 295, top 52% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1145 | #89 of 115, top 78% | EQ-Bench | ||
| LMArena Multi-Turn | 1339 | #165 of 295, top 56% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $2.50 | $10 | — | 2026-10-10 |
| cohere | $2.50 | $10 | — | 2026-10-10 |
| openrouter | $2.50 | $10 | — | 2026-10-10 |
Compare Command A
- Command A vs Granite 4.0 H Small
- Command A vs Grok Build 0.1
- Command A vs Llama 3.1 Nemotron Ultra 253b v1
- Command A vs gpt-oss-120b
- Command A vs o3-mini
- Command A vs Mistral Medium
- Command A vs GPT-6 Astra
- Command A vs Claude Fable 5.1
- Command A vs Gemini 3.8 Flash
- Command A vs Kimi K3
- Command A vs Grok 4.6
- Command A vs Qwen3.8 Max
- Command A vs GLM-5.3
- Command A vs Muse Spark 1.3
Other Cohere models
Frequently asked questions
How good is Command A?
Command A by Cohere ranks 215th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.5. Its strongest category is agentic & tool use, where it ranks 40th. API pricing starts at $2.50 per million input tokens and $10 per million output tokens, with a 256K-token context window.
How much does Command A cost?
Command A costs $2.50 per million input tokens and $10 per million output tokens on Cohere's own API.
What is Command A's context window?
Command A accepts up to 256K tokens of input and can write up to 8K tokens in one response.
Is Command A open source?
Yes. Command A's weights are downloadable from Hugging Face (CohereForAI/c4ai-command-a-03-2025); check the license for commercial terms.
How fast is Command A?
Command A generated about 28 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Command A's strengths and weaknesses?
Relative to other ranked models, Command A places best in agentic & tool use, knowledge, long context and lowest in coding, reasoning, writing & preference.
What is Command A best at?
Its best category is agentic & tool use, where it ranks 40th on Noometry.