Model comparison

Command R vs Muse Spark 1.2

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 31.4 on the Noometry Index. Command R costs 7.6× less per token, which makes it the better buy when Muse Spark 1.2's lead doesn't matter for your workload.

Last verified . 16 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Command R scores higher in 0 categories and Muse Spark 1.2 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Muse Spark 1.2 leads 51.3 to 13.8.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 94.7% for Muse Spark 1.2.
  • Command R is cheaper at $0.15 / $0.60 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.2.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 128K.
  • Command R has downloadable open weights; the other is API-only.

Side by side

Command R and Muse Spark 1.2 specifications
Command RMuse Spark 1.2
ProviderCohereMeta
Noometry Index31.450.3
Released2024-08-302026-08-05
WeightsOpenProprietary
Context window128K1.05M
Max output4K131K
Input $ / M tokens$0.15$1.25
Output $ / M tokens$0.60$4.25
Results tracked2931

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.2 leads

Command R: 29.3 (#306), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Coding11691495
DeepSWE—54.9%
LMArena WebDev—1533
FrontierSWE—12%
SciCode—56.4%
WeirdML—60.3%
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Agentic & Tool Use Not comparable

Command R: —, Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkCommand RMuse Spark 1.2
APEX-Agents—36.4%
GDP.pdf—16%

Reasoning Muse Spark 1.2 leads

Command R: 13.8 (#331), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Hard Prompts11641486
DTBench46.4%94.7%
LMCA9.2%48.4%
SimpleBench—74.5%
NYT Connections (extended)—79.2%
CritPt—17.7%
LiveBench Reasoning21.9%—
LiveBench Data Analysis33.3%—
Epoch Capabilities Index—154.87
LiveBench27.5%—

Math Muse Spark 1.2 leads

Command R: 28.0 (#246), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Math11551471
ProofBench—43%
LiveBench Math19.4%—

Knowledge Muse Spark 1.2 leads

Command R: 31.0 (#221), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Expert11381480
SimpleQA Verified—60.3%
MMLU65.2%—

Multimodal Not comparable

Command R: —, Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Vision—1305

Multilingual Muse Spark 1.2 leads

Command R: 35.7 (#245), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Non-English11741478
LMArena Chinese11821511
LMArena French11621513
LMArena Russian11741487
LMArena Spanish11511498
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—

Instruction Following Muse Spark 1.2 leads

Command R: 58.1 (#261), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Instruction Following11671461
LiveBench Instruction Following55.6%—

Long Context Muse Spark 1.2 leads

Command R: 36.3 (#231), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Longer Query11981475

Writing & Preference Muse Spark 1.2 leads

Command R: 38.2 (#254), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkCommand RMuse Spark 1.2
LMArena Text11871482
LMArena Creative Writing11701449
LMArena Multi-Turn11631494
EQ-Bench Creative Writing—1840
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Muse Spark 1.2?

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 31.4 on the Noometry Index. Command R costs 7.6× less per token, which makes it the better buy when Muse Spark 1.2's lead doesn't matter for your workload.

Which is cheaper, Command R or Muse Spark 1.2?

Command R is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Muse Spark 1.2 lists at $1.25 and $4.25.

Is Command R or Muse Spark 1.2 better for coding?

Muse Spark 1.2 scores higher on coding benchmarks: 49.2 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 128K.

How many benchmarks do Command R and Muse Spark 1.2 share?

16 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper