Model comparison

Command R+ vs GPT-4.1 nano

Command R+ is the stronger model overall, scoring 32.4 to 27.9 on the Noometry Index. GPT-4.1 nano costs 25× less per token, which makes it the better buy when Command R+'s lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Command R+ Cohere

32.4

Rank #257 Confirmed

GPT-4.1 nano OpenAI

27.9

Rank #327 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Command R+ scores higher in 6 categories and GPT-4.1 nano in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Command R+ leads 36.4 to 21.8.
  • GPT-4.1 nano is cheaper at $0.10 / $0.40 per million input/output tokens, against $2.50 / $10 for Command R+.
  • GPT-4.1 nano accepts more context: 1.05M tokens versus 128K.
  • Command R+ has downloadable open weights; the other is API-only.

Side by side

Command R+ and GPT-4.1 nano specifications
Command R+GPT-4.1 nano
ProviderCohereOpenAI
Noometry Index32.427.9
Released2024-08-302025-04-14
WeightsOpenProprietary
Context window128K1.05M
Max output4K33K
Input $ / M tokens$2.50$0.10
Output $ / M tokens$10$0.40
Results tracked3438

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Command R+ leads

Command R+: 29.1 (#309), GPT-4.1 nano: 24.1 (#330)

Coding benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Coding11871306
Aider Polyglot—8.9%
SciCode—25.9%
WeirdML—19%
BigCodeBench Instruct33.8%—
LiveBench Coding19.1%—
BigCodeBench Complete41.9%—
HumanEval+56.7%—
MBPP+63.5%—

Agentic & Tool Use Not comparable

Command R+: —, GPT-4.1 nano: 26.5 (#104)

Agentic & Tool Use benchmarks
BenchmarkCommand R+GPT-4.1 nano
Berkeley Function Calling Leaderboard—33%

Reasoning Too close to call

Command R+: 9.2 (#344), GPT-4.1 nano: 8.5 (#349)

Reasoning benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Hard Prompts11861286
DTBench54.9%52.5%
LMCA5%5.5%
Epoch Capabilities Index119.34129.62
ARC-AGI-2—0%
SimpleBench17.4%—
Kagi LLM Benchmark—33.3%
ARC-AGI-1—0%
CritPt—0%
LiveBench Reasoning24.8%—
LiveBench Data Analysis38.1%—
LiveBench31.8%—

Math Command R+ leads

Command R+: 28.9 (#242), GPT-4.1 nano: 26.9 (#252)

Math benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Math11881274
OTIS Mock AIME 2024-2025—28.9%
Omni-MATH—36.7%
LiveBench Math21.3%—
MATH Level 5—70%
FrontierMath (Feb 2025 set)—1%

Knowledge Command R+ leads

Command R+: 36.4 (#169), GPT-4.1 nano: 21.8 (#273)

Knowledge benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Expert11741272
GPQA Diamond—48.9%
SimpleQA Verified—6%
MMLU-Pro—55%
Vectara Hallucination Rate6.9%—
GPQA (HELM)—50.7%
MMLU69.4%—

Multimodal Not comparable

Command R+: —, GPT-4.1 nano: 29.2 (#113)

Multimodal benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Vision—1063

Multilingual GPT-4.1 nano leads

Command R+: 38.6 (#227), GPT-4.1 nano: 41.6 (#205)

Multilingual benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Non-English12161260
LMArena Chinese12261270
LMArena German12161288
LMArena Japanese11661198
LMArena Russian12271261
LMArena French1209—
LMArena Korean1138—
LMArena Spanish1189—

Instruction Following GPT-4.1 nano leads

Command R+: 60.0 (#254), GPT-4.1 nano: 67.8 (#193)

Instruction Following benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Instruction Following11971267
LiveBench Instruction Following57.6%—
IFEval—84.3%

Long Context Command R+ leads

Command R+: 37.3 (#219), GPT-4.1 nano: 23.7 (#296)

Long Context benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Longer Query12301283
Fiction.LiveBench—25%

Writing & Preference Command R+ leads

Command R+: 43.5 (#228), GPT-4.1 nano: 40.5 (#243)

Writing & Preference benchmarks
BenchmarkCommand R+GPT-4.1 nano
LMArena Text12291285
LMArena Creative Writing12351260
LMArena Multi-Turn12131277
EQ-Bench Creative Writing—946
WildBench—81.2%
LiveBench Language29.7%—

Frequently asked questions

Is Command R+ better than GPT-4.1 nano?

Command R+ is the stronger model overall, scoring 32.4 to 27.9 on the Noometry Index. GPT-4.1 nano costs 25× less per token, which makes it the better buy when Command R+'s lead doesn't matter for your workload.

Which is cheaper, Command R+ or GPT-4.1 nano?

GPT-4.1 nano is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Command R+ lists at $2.50 and $10.

Is Command R+ or GPT-4.1 nano better for coding?

Command R+ scores higher on coding benchmarks: 29.1 versus 24.1 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 nano does, with 1.05M tokens against 128K.

How many benchmarks do Command R+ and GPT-4.1 nano share?

17 benchmarks have published results for both models. Command R+ has 34 scored results on Noometry and GPT-4.1 nano has 38.

Related comparisons

Go deeper