Model comparison

Command R+ vs Phi-4

Command R+ is the stronger model overall, scoring 32.4 to 31.2 on the Noometry Index. Phi-4 costs 50× less per token, which makes it the better buy when Command R+'s lead doesn't matter for your workload.

Last verified . 29 shared benchmarks.

Command R+ Cohere

32.4

Rank #257 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Command R+ scores higher in 5 categories and Phi-4 in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Phi-4 leads 17.7 to 9.2.
  • The biggest single-benchmark swing is LiveBench Reasoning: 24.8% for Command R+ and 47.8% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $2.50 / $10 for Command R+.

Side by side

Command R+ and Phi-4 specifications
Command R+Phi-4
ProviderCohereMicrosoft
Noometry Index32.431.2
Released2024-08-302024-12-11
WeightsOpenOpen
Context window128K128K
Max output4K4K
Input $ / M tokens$2.50$0.07
Output $ / M tokens$10$0.14
Results tracked3437

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi-4 leads

Command R+: 29.1 (#309), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkCommand R+Phi-4
BigCodeBench Instruct33.8%45.5%
LiveBench Coding19.1%30.7%
LMArena Coding11871231
BigCodeBench Complete41.9%55.4%
HumanEval+56.7%—
MBPP+63.5%—

Agentic & Tool Use Not comparable

Command R+: —, Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkCommand R+Phi-4
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%

Reasoning Phi-4 leads

Command R+: 9.2 (#344), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkCommand R+Phi-4
LiveBench Reasoning24.8%47.8%
LMArena Hard Prompts11861220
LiveBench Data Analysis38.1%45.2%
Epoch Capabilities Index119.34130.42
LiveBench31.8%41.6%
SimpleBench17.4%—
Chess Puzzles—1%
DTBench54.9%—
LMCA5%—

Math Command R+ leads

Command R+: 28.9 (#242), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkCommand R+Phi-4
LiveBench Math21.3%42%
LMArena Math11881246
OTIS Mock AIME 2024-2025—13.8%
MATH Level 5—64.9%

Knowledge Command R+ leads

Command R+: 36.4 (#169), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkCommand R+Phi-4
Vectara Hallucination Rate6.9%3.7%
LMArena Expert11741203
MMLU69.4%84.8%
GPQA Diamond—56.1%
Confabulations—29.4%

Multilingual Command R+ leads

Command R+: 38.6 (#227), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkCommand R+Phi-4
LMArena Non-English12161197
LMArena Chinese12261212
LMArena French12091224
LMArena German12161222
LMArena Japanese11661158
LMArena Korean11381151
LMArena Russian12271209
LMArena Spanish11891234

Instruction Following Too close to call

Command R+: 60.0 (#254), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkCommand R+Phi-4
LiveBench Instruction Following57.6%58.4%
LMArena Instruction Following11971201

Long Context Too close to call

Command R+: 37.3 (#219), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkCommand R+Phi-4
LMArena Longer Query12301217

Writing & Preference Command R+ leads

Command R+: 43.5 (#228), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkCommand R+Phi-4
LMArena Text12291217
LMArena Creative Writing12351182
LMArena Multi-Turn12131206
LiveBench Language29.7%25.6%
Short-Story Creative Writing—62.6%

Frequently asked questions

Is Command R+ better than Phi-4?

Command R+ is the stronger model overall, scoring 32.4 to 31.2 on the Noometry Index. Phi-4 costs 50× less per token, which makes it the better buy when Command R+'s lead doesn't matter for your workload.

Which is cheaper, Command R+ or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Command R+ lists at $2.50 and $10.

Is Command R+ or Phi-4 better for coding?

Phi-4 scores higher on coding benchmarks: 34.4 versus 29.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Command R+ and Phi-4 share?

29 benchmarks have published results for both models. Command R+ has 34 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper