Model comparison

Command R vs Mistral Small 3.1

Command R and Mistral Small 3.1 score almost the same on the Noometry Index (31.4 vs 31.7), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Command R scores higher in 3 categories and Mistral Small 3.1 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Command R leads 28.0 to 14.7.
  • Command R is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.

Side by side

Command R and Mistral Small 3.1 specifications
Command RMistral Small 3.1
ProviderCohereMistral AI
Noometry Index31.431.7
Released2024-08-302025-03-17
WeightsOpenOpen
Context window128K128K
Max output4K102K
Input $ / M tokens$0.15$0.35
Output $ / M tokens$0.60$0.56
Results tracked2928

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Command R: 29.3 (#306), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Coding11691309
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Reasoning Mistral Small 3.1 leads

Command R: 13.8 (#331), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Hard Prompts11641278
Chess Puzzles—1%
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
Epoch Capabilities Index—127.48
LiveBench27.5%—

Math Command R leads

Command R: 28.0 (#246), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Math11551262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%
LiveBench Math19.4%—

Knowledge Command R leads

Command R: 31.0 (#221), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Expert11381257
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%
MMLU65.2%—

Multimodal Not comparable

Command R: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Vision—1136

Multilingual Mistral Small 3.1 leads

Command R: 35.7 (#245), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Non-English11741255
LMArena Chinese11821253
LMArena French11621273
LMArena German11761266
LMArena Japanese11431208
LMArena Korean11631206
LMArena Russian11741263
LMArena Spanish11511283

Instruction Following Mistral Small 3.1 leads

Command R: 58.1 (#261), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Instruction Following11671264
LiveBench Instruction Following55.6%—
IFEval—75%

Long Context Mistral Small 3.1 leads

Command R: 36.3 (#231), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Longer Query11981299

Writing & Preference Command R leads

Command R: 38.2 (#254), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkCommand RMistral Small 3.1
LMArena Text11871277
LMArena Creative Writing11701253
LMArena Multi-Turn11631270
EQ-Bench Creative Writing—761
WildBench—78.8%
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Mistral Small 3.1?

Command R and Mistral Small 3.1 score almost the same on the Noometry Index (31.4 vs 31.7), so choose on price, context window or the category you care about most.

Which is cheaper, Command R or Mistral Small 3.1?

Command R is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Command R or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Command R and Mistral Small 3.1 share?

17 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper