Model comparison

Command A vs Qwen1.5-110B

Command A is the stronger model overall, scoring 36.5 to 34.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Command A scores higher in 6 categories and Qwen1.5-110B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Command A leads 45.3 to 33.6.

Side by side

Command A and Qwen1.5-110B specifications
Command AQwen1.5-110B
ProviderCohereAlibaba (Qwen)
Noometry Index36.534.2
Released2025-03-132024-04-25
WeightsOpenOpen
Context window256K—
Max output8K—
Input $ / M tokens$2.50—
Output $ / M tokens$10—
Results tracked2420

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-110B leads

Command A: 27.2 (#322), Qwen1.5-110B: 33.0 (#264)

Coding benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Coding13301184
Aider Polyglot12%—
BigCodeBench Instruct—35%
BigCodeBench Complete—44.4%

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Qwen1.5-110B: —

Agentic & Tool Use benchmarks
BenchmarkCommand AQwen1.5-110B
Berkeley Function Calling Leaderboard57.1%—

Reasoning Qwen1.5-110B leads

Command A: 18.3 (#283), Qwen1.5-110B: 22.7 (#189)

Reasoning benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Hard Prompts13261168
Kagi LLM Benchmark28.8%—
DTBench61.3%—
LMCA10.3%—
ForecastBench—57.7

Math Command A leads

Command A: 36.2 (#171), Qwen1.5-110B: 33.7 (#201)

Math benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Math13001185

Knowledge Command A leads

Command A: 37.1 (#159), Qwen1.5-110B: 31.2 (#219)

Knowledge benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Expert12951144
Vectara Hallucination Rate9.3%—

Multilingual Command A leads

Command A: 45.3 (#170), Qwen1.5-110B: 33.6 (#250)

Multilingual benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Non-English13131142
LMArena Chinese13271206
LMArena French13511151
LMArena German13411123
LMArena Japanese12851074
LMArena Korean12851044
LMArena Russian13141118
LMArena Spanish13471142

Instruction Following Command A leads

Command A: 69.1 (#177), Qwen1.5-110B: 60.3 (#252)

Instruction Following benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Instruction Following13091158

Long Context Command A leads

Command A: 40.6 (#151), Qwen1.5-110B: 35.1 (#242)

Long Context benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Longer Query13341157

Writing & Preference Command A leads

Command A: 47.6 (#208), Qwen1.5-110B: 38.0 (#255)

Writing & Preference benchmarks
BenchmarkCommand AQwen1.5-110B
LMArena Text13311175
LMArena Creative Writing13191148
LMArena Multi-Turn13391160
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Qwen1.5-110B?

Command A is the stronger model overall, scoring 36.5 to 34.2 on the Noometry Index.

Is Command A or Qwen1.5-110B better for coding?

Qwen1.5-110B scores higher on coding benchmarks: 33.0 versus 27.2 in the Noometry coding category.

How many benchmarks do Command A and Qwen1.5-110B share?

17 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Qwen1.5-110B has 20.

Related comparisons

Go deeper