Model comparison

Command A vs Gemini 2.0 Flash (Feb 2025)

Command A is the stronger model overall, scoring 36.5 to 35.1 on the Noometry Index.

Last verified . 21 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Command A scores higher in 4 categories and Gemini 2.0 Flash (Feb 2025) in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Command A leads 35.9 to 28.1.
  • The biggest single-benchmark swing is Aider Polyglot: 12% for Command A and 38.2% for Gemini 2.0 Flash (Feb 2025).
  • Command A has downloadable open weights; the other is API-only.

Side by side

Command A and Gemini 2.0 Flash (Feb 2025) specifications
Command AGemini 2.0 Flash (Feb 2025)
ProviderCohereGoogle
Noometry Index36.535.1
Released2025-03-132024-12-06
WeightsOpenProprietary
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2454

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.0 Flash (Feb 2025) leads

Command A: 27.2 (#322), Gemini 2.0 Flash (Feb 2025): 28.4 (#315)

Coding benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
Aider Polyglot12%38.2%
LMArena Coding13301350
SWE-bench Verified (bash only)—13.5%
WeirdML—25.8%
BigCodeBench Instruct—45.9%
LiveBench Coding—63.4%
BigCodeBench Complete—59.9%
CadEval—30%

Agentic & Tool Use Command A leads

Command A: 35.9 (#40), Gemini 2.0 Flash (Feb 2025): 28.1 (#92)

Agentic & Tool Use benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
Berkeley Function Calling Leaderboard57.1%—
TheAgentCompany—11.4%

Reasoning Command A leads

Command A: 18.3 (#283), Gemini 2.0 Flash (Feb 2025): 15.2 (#318)

Reasoning benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
Kagi LLM Benchmark28.8%37.8%
LMArena Hard Prompts13261346
DTBench61.3%63.2%
ARC-AGI-2—1.3%
SimpleBench—31.1%
EnigmaEval—1.1%
LiveBench Reasoning—78.2%
LiveBench Data Analysis—69.4%
LMCA10.3%—
Epoch Capabilities Index—135.36
LiveBench—66.9%

Math Gemini 2.0 Flash (Feb 2025) leads

Command A: 36.2 (#171), Gemini 2.0 Flash (Feb 2025): 37.9 (#146)

Math benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Math13001352
OTIS Mock AIME 2024-2025—57.8%
Omni-MATH—45.9%
LiveBench Math—75.8%
MATH Level 5—82.2%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Command A leads

Command A: 37.1 (#159), Gemini 2.0 Flash (Feb 2025): 32.0 (#213)

Knowledge benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Expert12951339
GPQA Diamond—64.1%
Humanity's Last Exam—6.6%
MMLU-Pro—73.7%
Confabulations—12.4%
Vectara Hallucination Rate9.3%—
GPQA (HELM)—55.6%
MMLU—79.7%

Multimodal Not comparable

Command A: —, Gemini 2.0 Flash (Feb 2025): 36.5 (#79)

Multimodal benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Vision—1158
GeoBench—77%

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Command A: 45.3 (#170), Gemini 2.0 Flash (Feb 2025): 47.4 (#149)

Multilingual benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Non-English13131342
LMArena Chinese13271373
LMArena French13511391
LMArena German13411353
LMArena Japanese12851294
LMArena Korean12851313
LMArena Russian13141351
LMArena Spanish13471363

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Command A: 69.1 (#177), Gemini 2.0 Flash (Feb 2025): 74.4 (#97)

Instruction Following benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Instruction Following13091336
LiveBench Instruction Following—85.8%
IFEval—84.1%

Long Context Command A leads

Command A: 40.6 (#151), Gemini 2.0 Flash (Feb 2025): 38.1 (#203)

Long Context benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Longer Query13341344
Fiction.LiveBench—61.1%

Writing & Preference Gemini 2.0 Flash (Feb 2025) leads

Command A: 47.6 (#208), Gemini 2.0 Flash (Feb 2025): 49.5 (#190)

Writing & Preference benchmarks
BenchmarkCommand AGemini 2.0 Flash (Feb 2025)
LMArena Text13311354
LMArena Creative Writing13191340
EQ-Bench Creative Writing11451128
LMArena Multi-Turn13391350
Short-Story Creative Writing—73.8%
WildBench—80%
LiveBench Language—51.3%

Frequently asked questions

Is Command A better than Gemini 2.0 Flash (Feb 2025)?

Command A is the stronger model overall, scoring 36.5 to 35.1 on the Noometry Index.

Is Command A or Gemini 2.0 Flash (Feb 2025) better for coding?

Gemini 2.0 Flash (Feb 2025) scores higher on coding benchmarks: 28.4 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Gemini 2.0 Flash (Feb 2025) share?

21 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Gemini 2.0 Flash (Feb 2025) has 54.

Related comparisons

Go deeper