Model comparison

Command R vs Phi-4

Command R and Phi-4 score almost the same on the Noometry Index (31.4 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Command R scores higher in 1 category and Phi-4 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Command R leads 28.0 to 20.8.
  • The biggest single-benchmark swing is LiveBench Reasoning: 21.9% for Command R and 47.8% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $0.15 / $0.60 for Command R.

Side by side

Command R and Phi-4 specifications
Command RPhi-4
ProviderCohereMicrosoft
Noometry Index31.431.2
Released2024-08-302024-12-11
WeightsOpenOpen
Context window128K128K
Max output4K4K
Input $ / M tokens$0.15$0.07
Output $ / M tokens$0.60$0.14
Results tracked2937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi-4 leads

Command R: 29.3 (#306), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkCommand RPhi-4
BigCodeBench Instruct37.1%45.5%
LiveBench Coding17.9%30.7%
LMArena Coding11691231
BigCodeBench Complete45.2%55.4%

Agentic & Tool Use Not comparable

Command R: —, Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkCommand RPhi-4
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%

Reasoning Phi-4 leads

Command R: 13.8 (#331), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkCommand RPhi-4
LiveBench Reasoning21.9%47.8%
LMArena Hard Prompts11641220
LiveBench Data Analysis33.3%45.2%
LiveBench27.5%41.6%
Chess Puzzles—1%
DTBench46.4%—
LMCA9.2%—
Epoch Capabilities Index—130.42

Math Command R leads

Command R: 28.0 (#246), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkCommand RPhi-4
LiveBench Math19.4%42%
LMArena Math11551246
OTIS Mock AIME 2024-2025—13.8%
MATH Level 5—64.9%

Knowledge Phi-4 leads

Command R: 31.0 (#221), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkCommand RPhi-4
LMArena Expert11381203
MMLU65.2%84.8%
GPQA Diamond—56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%

Multilingual Phi-4 leads

Command R: 35.7 (#245), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkCommand RPhi-4
LMArena Non-English11741197
LMArena Chinese11821212
LMArena French11621224
LMArena German11761222
LMArena Japanese11431158
LMArena Korean11631151
LMArena Russian11741209
LMArena Spanish11511234

Instruction Following Phi-4 leads

Command R: 58.1 (#261), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkCommand RPhi-4
LiveBench Instruction Following55.6%58.4%
LMArena Instruction Following11671201

Long Context Too close to call

Command R: 36.3 (#231), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkCommand RPhi-4
LMArena Longer Query11981217

Writing & Preference Phi-4 leads

Command R: 38.2 (#254), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkCommand RPhi-4
LMArena Text11871217
LMArena Creative Writing11701182
LMArena Multi-Turn11631206
LiveBench Language16.7%25.6%
Short-Story Creative Writing—62.6%

Frequently asked questions

Is Command R better than Phi-4?

Command R and Phi-4 score almost the same on the Noometry Index (31.4 vs 31.2), so choose on price, context window or the category you care about most.

Which is cheaper, Command R or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Command R lists at $0.15 and $0.60.

Is Command R or Phi-4 better for coding?

Phi-4 scores higher on coding benchmarks: 34.4 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Command R and Phi-4 share?

27 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper