Model comparison

Command R vs gpt-oss-20b

gpt-oss-20b is the stronger model overall, scoring 32.5 to 31.4 on the Noometry Index.

Last verified . 18 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Command R scores higher in 1 category and gpt-oss-20b in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-20b leads 39.4 to 28.0.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 68% for gpt-oss-20b.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.15 / $0.60 for Command R.
  • gpt-oss-20b accepts more context: 131K tokens versus 128K.

Side by side

Command R and gpt-oss-20b specifications
Command Rgpt-oss-20b
ProviderCohereOpenAI
Noometry Index31.432.5
Released2024-08-302025-08-05
WeightsOpenOpen
Context window128K131K
Max output4K16K
Input $ / M tokens$0.15$0.018
Output $ / M tokens$0.60$0.09
Results tracked2934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding gpt-oss-20b leads

Command R: 29.3 (#306), gpt-oss-20b: 37.6 (#192)

Coding benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Coding11691306
SciCode—34.4%
WeirdML—40.9%
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—
ALE-Bench—566.05

Agentic & Tool Use Not comparable

Command R: —, gpt-oss-20b: 9.3 (#154)

Agentic & Tool Use benchmarks
BenchmarkCommand Rgpt-oss-20b
Terminal-Bench—3.4%

Reasoning gpt-oss-20b leads

Command R: 13.8 (#331), gpt-oss-20b: 19.3 (#261)

Reasoning benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Hard Prompts11641274
DTBench46.4%68%
LMCA9.2%14.5%
Kagi LLM Benchmark—53.2%
CritPt—1.4%
Chess Puzzles—4%
LiveBench Reasoning21.9%—
LiveBench Data Analysis33.3%—
Epoch Capabilities Index—137.82
LiveBench27.5%—

Math gpt-oss-20b leads

Command R: 28.0 (#246), gpt-oss-20b: 39.4 (#103)

Math benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Math11551317
OTIS Mock AIME 2024-2025—65.3%
Omni-MATH—56.5%
LiveBench Math19.4%—

Knowledge gpt-oss-20b leads

Command R: 31.0 (#221), gpt-oss-20b: 34.6 (#195)

Knowledge benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Expert11381258
GPQA Diamond—60.8%
MMLU-Pro—74%
GPQA (HELM)—59.4%
MMLU65.2%—

Multilingual gpt-oss-20b leads

Command R: 35.7 (#245), gpt-oss-20b: 42.2 (#197)

Multilingual benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Non-English11741268
LMArena Chinese11821314
LMArena German11761255
LMArena Japanese11431244
LMArena Korean11631236
LMArena Russian11741278
LMArena Spanish11511267
LMArena French1162—

Instruction Following gpt-oss-20b leads

Command R: 58.1 (#261), gpt-oss-20b: 61.8 (#240)

Instruction Following benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Instruction Following11671236
LiveBench Instruction Following55.6%—
IFEval—73.2%

Long Context gpt-oss-20b leads

Command R: 36.3 (#231), gpt-oss-20b: 37.9 (#209)

Long Context benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Longer Query11981250

Writing & Preference Command R leads

Command R: 38.2 (#254), gpt-oss-20b: 35.5 (#265)

Writing & Preference benchmarks
BenchmarkCommand Rgpt-oss-20b
LMArena Text11871287
LMArena Creative Writing11701201
LMArena Multi-Turn11631268
EQ-Bench Creative Writing—666
WildBench—73.7%
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than gpt-oss-20b?

gpt-oss-20b is the stronger model overall, scoring 32.5 to 31.4 on the Noometry Index.

Which is cheaper, Command R or gpt-oss-20b?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Command R lists at $0.15 and $0.60.

Is Command R or gpt-oss-20b better for coding?

gpt-oss-20b scores higher on coding benchmarks: 37.6 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

gpt-oss-20b does, with 131K tokens against 128K.

How many benchmarks do Command R and gpt-oss-20b share?

18 benchmarks have published results for both models. Command R has 29 scored results on Noometry and gpt-oss-20b has 34.

Related comparisons

Go deeper