Model comparison

Command R vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 31.4 on the Noometry Index. Command R costs 9.1× less per token, which makes it the better buy when Qwen3 Max's lead doesn't matter for your workload.

Last verified . 19 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Command R scores higher in 0 categories and Qwen3 Max in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3 Max leads 62.4 to 38.2.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 82.1% for Qwen3 Max.
  • Command R is cheaper at $0.15 / $0.60 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Qwen3 Max accepts more context: 262K tokens versus 128K.
  • Command R has downloadable open weights; the other is API-only.

Side by side

Command R and Qwen3 Max specifications
Command RQwen3 Max
ProviderCohereAlibaba (Qwen)
Noometry Index31.443.7
Released2024-08-302025-09-23
WeightsOpenProprietary
Context window128K262K
Max output4K66K
Input $ / M tokens$0.15$1.20
Output $ / M tokens$0.60$6
Results tracked2933

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 Max leads

Command R: 29.3 (#306), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkCommand RQwen3 Max
LMArena Coding11691456
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—
ALE-Bench—370.45

Agentic & Tool Use Not comparable

Command R: —, Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkCommand RQwen3 Max
Vending-Bench 2—71.56

Reasoning Qwen3 Max leads

Command R: 13.8 (#331), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkCommand RQwen3 Max
LMArena Hard Prompts11641448
DTBench46.4%82.1%
LMCA9.2%28.3%
Kagi LLM Benchmark—72.5%
NYT Connections (extended)—30.1%
Chess Puzzles—4%
LiveBench Reasoning21.9%—
Mystery Game Puzzles—5%
LiveBench Data Analysis33.3%—
Epoch Capabilities Index—142.38
LiveBench27.5%—

Math Qwen3 Max leads

Command R: 28.0 (#246), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkCommand RQwen3 Max
LMArena Math11551446
FrontierMath (Tiers 1-3)—18.9%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math19.4%—
MATH Level 5—97.1%

Knowledge Qwen3 Max leads

Command R: 31.0 (#221), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkCommand RQwen3 Max
LMArena Expert11381455
GPQA Diamond—72.6%
SimpleQA Verified—48.7%
MMLU65.2%—

Multilingual Qwen3 Max leads

Command R: 35.7 (#245), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkCommand RQwen3 Max
LMArena Non-English11741429
LMArena Chinese11821478
LMArena French11621449
LMArena German11761463
LMArena Japanese11431397
LMArena Korean11631399
LMArena Russian11741428
LMArena Spanish11511462

Instruction Following Qwen3 Max leads

Command R: 58.1 (#261), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkCommand RQwen3 Max
LMArena Instruction Following11671419
LiveBench Instruction Following55.6%—

Long Context Qwen3 Max leads

Command R: 36.3 (#231), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkCommand RQwen3 Max
LMArena Longer Query11981438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Command R: 38.2 (#254), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkCommand RQwen3 Max
LMArena Text11871439
LMArena Creative Writing11701402
LMArena Multi-Turn11631446
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 31.4 on the Noometry Index. Command R costs 9.1× less per token, which makes it the better buy when Qwen3 Max's lead doesn't matter for your workload.

Which is cheaper, Command R or Qwen3 Max?

Command R is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Command R or Qwen3 Max better for coding?

Qwen3 Max scores higher on coding benchmarks: 43.0 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

Qwen3 Max does, with 262K tokens against 128K.

How many benchmarks do Command R and Qwen3 Max share?

19 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper