Model comparison

DeepSeek-R1 vs Mistral Large

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 31.9 on the Noometry Index.

Last verified . 42 shared benchmarks.

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 42 benchmarks with published results for both. DeepSeek-R1 scores higher in 9 categories and Mistral Large in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-R1 leads 43.8 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 66.4% for DeepSeek-R1 and 8.5% for Mistral Large.
  • DeepSeek-R1 is cheaper at $0.50 / $2.15 per million input/output tokens, against $2 / $6 for Mistral Large.
  • DeepSeek-R1 accepts more context: 164K tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1 and Mistral Large specifications
DeepSeek-R1Mistral Large
ProviderDeepSeekMistral AI
Noometry Index42.331.9
Released2025-01-202024-02-26
WeightsProprietaryOpen
Context window164K131K
Max output64K16K
Input $ / M tokens$0.50$2
Output $ / M tokens$2.15$6
Results tracked5251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

DeepSeek-R1: 46.3 (#68), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkDeepSeek-R1Mistral Large
SciCode35.7%36.2%
LiveBench Coding66.7%47.1%
LMArena Coding14271277
ALE-Bench804.12264.7
Aider Polyglot71.4%—
WeirdML41.6%—
BigCodeBench Instruct—30%
BigCodeBench Complete—38.3%
AlgoTune1.7—
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use DeepSeek-R1 leads

DeepSeek-R1: 30.7 (#75), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1Mistral Large
Berkeley Function Calling Leaderboard—38.4%
DeepResearch Bench35.1%—
BALROG34.9%—
METR Time Horizons53.8%—

Reasoning DeepSeek-R1 leads

DeepSeek-R1: 18.6 (#278), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkDeepSeek-R1Mistral Large
SimpleBench40.8%22.5%
CritPt1.1%0%
LiveBench Reasoning83.2%43.5%
LMArena Hard Prompts14161257
LiveBench Data Analysis69.8%50.1%
Epoch Capabilities Index141.29128.52
ForecastBench6057.1
LiveBench71.6%48.4%
ARC-AGI-21.3%—
Kagi LLM Benchmark69.4%—
ARC-AGI-121.2%—
DTBench—65.1%
LMCA—16.7%

Math DeepSeek-R1 leads

DeepSeek-R1: 43.8 (#79), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkDeepSeek-R1Mistral Large
OTIS Mock AIME 2024-202566.4%8.5%
Omni-MATH42.4%28.1%
LiveBench Math80.7%42.5%
LMArena Math14001262
MATH Level 596.6%50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge DeepSeek-R1 leads

DeepSeek-R1: 44.5 (#87), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkDeepSeek-R1Mistral Large
GPQA Diamond76.3%51.3%
MMLU-Pro79.3%59.9%
Confabulations12.7%21.4%
Vectara Hallucination Rate11.3%4.5%
GPQA (HELM)66.6%43.5%
LMArena Expert13941232
MMLU—80%

Multilingual DeepSeek-R1 leads

DeepSeek-R1: 52.4 (#85), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkDeepSeek-R1Mistral Large
LMArena Non-English14121237
LMArena Chinese14421240
LMArena French14171325
LMArena German14041254
LMArena Japanese13911188
LMArena Korean13601202
LMArena Russian14231257
LMArena Spanish14111268

Instruction Following DeepSeek-R1 leads

DeepSeek-R1: 72.0 (#143), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkDeepSeek-R1Mistral Large
LiveBench Instruction Following80.5%67.9%
IFEval78.4%87.7%
LMArena Instruction Following13821249

Long Context DeepSeek-R1 leads

DeepSeek-R1: 45.4 (#36), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkDeepSeek-R1Mistral Large
LMArena Longer Query13911261
Fiction.LiveBench75%—

Writing & Preference DeepSeek-R1 leads

DeepSeek-R1: 61.4 (#88), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1Mistral Large
LMArena Text14281266
LMArena Creative Writing14051243
Short-Story Creative Writing83%69%
EQ-Bench Creative Writing1500985
WildBench82.8%80.1%
LMArena Multi-Turn14051260
LiveBench Language48.5%39.4%

Frequently asked questions

Is DeepSeek-R1 better than Mistral Large?

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 31.9 on the Noometry Index.

Which is cheaper, DeepSeek-R1 or Mistral Large?

DeepSeek-R1 is cheaper. It lists at $0.50 per million input tokens and $2.15 per million output tokens; Mistral Large lists at $2 and $6.

Is DeepSeek-R1 or Mistral Large better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-R1 does, with 164K tokens against 131K.

How many benchmarks do DeepSeek-R1 and Mistral Large share?

42 benchmarks have published results for both models. DeepSeek-R1 has 52 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper