Model comparison

Command R vs GPT-6 Luna

GPT-6 Luna is the stronger model overall, scoring 53.3 to 31.4 on the Noometry Index.

Last verified . 19 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

GPT-6 Luna OpenAI

53.3

Rank #36 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Command R scores higher in 0 categories and GPT-6 Luna in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-6 Luna leads 76.1 to 28.0.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 90.1% for GPT-6 Luna.
  • GPT-6 Luna is cheaper at $0.10 / $0.50 per million input/output tokens, against $0.15 / $0.60 for Command R.
  • GPT-6 Luna accepts more context: 1.05M tokens versus 128K.
  • Command R has downloadable open weights; the other is API-only.

Side by side

Command R and GPT-6 Luna specifications
Command RGPT-6 Luna
ProviderCohereOpenAI
Noometry Index31.453.3
Released2024-08-302026-09-22
WeightsOpenProprietary
Context window128K1.05M
Max output4K128K
Input $ / M tokens$0.15$0.10
Output $ / M tokens$0.60$0.50
Results tracked2942

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Luna leads

Command R: 29.3 (#306), GPT-6 Luna: 55.5 (#25)

Coding benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Coding11691439
DeepSWE—66.6%
FrontierCode—42.4%
LMArena WebDev—1581
SciCode—54.6%
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—
ALE-Bench—1,577

Agentic & Tool Use Not comparable

Command R: —, GPT-6 Luna: 33.3 (#54)

Agentic & Tool Use benchmarks
BenchmarkCommand RGPT-6 Luna
APEX-Agents—44.3%
GDP.pdf—23%

Reasoning GPT-6 Luna leads

Command R: 13.8 (#331), GPT-6 Luna: 48.2 (#41)

Reasoning benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Hard Prompts11641411
DTBench46.4%90.1%
LMCA9.2%44.5%
ARC-AGI-2—59.3%
NYT Connections (extended)—68.7%
ARC-AGI-1—86.7%
CritPt—19.4%
Chess Puzzles—31%
LiveBench Reasoning21.9%—
Mystery Game Puzzles—7%
LiveBench Data Analysis33.3%—
Epoch Capabilities Index—156.28
LiveBench27.5%—

Math GPT-6 Luna leads

Command R: 28.0 (#246), GPT-6 Luna: 76.1 (#15)

Math benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Math11551416
FrontierMath (Tiers 1-3)—78.9%
FrontierMath Tier 4—56.1%
OTIS Mock AIME 2024-2025—98.9%
ProofBench—64%
LiveBench Math19.4%—

Knowledge GPT-6 Luna leads

Command R: 31.0 (#221), GPT-6 Luna: 57.0 (#41)

Knowledge benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Expert11381444
GPQA Diamond—90.5%
SimpleQA Verified—41.4%
MMLU65.2%—

Multimodal Not comparable

Command R: —, GPT-6 Luna: 42.4 (#30)

Multimodal benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Vision—1217
Blueprint-Bench 2—31.2%
Furniture Assembly—44.2%

Multilingual GPT-6 Luna leads

Command R: 35.7 (#245), GPT-6 Luna: 50.5 (#117)

Multilingual benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Non-English11741386
LMArena Chinese11821433
LMArena French11621420
LMArena German11761369
LMArena Japanese11431369
LMArena Korean11631360
LMArena Russian11741394
LMArena Spanish11511393

Instruction Following GPT-6 Luna leads

Command R: 58.1 (#261), GPT-6 Luna: 74.3 (#99)

Instruction Following benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Instruction Following11671409
LiveBench Instruction Following55.6%—

Long Context GPT-6 Luna leads

Command R: 36.3 (#231), GPT-6 Luna: 43.0 (#111)

Long Context benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Longer Query11981409

Writing & Preference GPT-6 Luna leads

Command R: 38.2 (#254), GPT-6 Luna: 58.3 (#119)

Writing & Preference benchmarks
BenchmarkCommand RGPT-6 Luna
LMArena Text11871391
LMArena Creative Writing11701363
LMArena Multi-Turn11631396
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than GPT-6 Luna?

GPT-6 Luna is the stronger model overall, scoring 53.3 to 31.4 on the Noometry Index.

Which is cheaper, Command R or GPT-6 Luna?

GPT-6 Luna is cheaper. It lists at $0.10 per million input tokens and $0.50 per million output tokens; Command R lists at $0.15 and $0.60.

Is Command R or GPT-6 Luna better for coding?

GPT-6 Luna scores higher on coding benchmarks: 55.5 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Luna does, with 1.05M tokens against 128K.

How many benchmarks do Command R and GPT-6 Luna share?

19 benchmarks have published results for both models. Command R has 29 scored results on Noometry and GPT-6 Luna has 42.

Related comparisons

Go deeper