Model comparison

DeepSeek-V3 vs Mistral Medium 3.5

DeepSeek-V3 and Mistral Medium 3.5 score almost the same on the Noometry Index (39.5 vs 40.2), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DeepSeek-V3 scores higher in 2 categories and Mistral Medium 3.5 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Mistral Medium 3.5 leads 43.2 to 34.0.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 52.3% for DeepSeek-V3 and 41.4% for Mistral Medium 3.5.
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 164K.

Side by side

DeepSeek-V3 and Mistral Medium 3.5 specifications
DeepSeek-V3Mistral Medium 3.5
ProviderDeepSeekMistral AI
Noometry Index39.540.2
Released2024-12-26—
WeightsOpenOpen
Context window164K262K
Max output164K210K
Input $ / M tokens$0.24$1.50
Output $ / M tokens$0.90$7.50
Results tracked6022

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Coding13681461
Aider Polyglot55.1%—
LMArena WebDev—1264
SciCode35.8%—
WeirdML36.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Mistral Medium 3.5: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
METR Time Horizons49.6%—

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
Kagi LLM Benchmark52.3%41.4%
LMArena Hard Prompts13651436
Epoch Capabilities Index135.94141.35
SimpleBench27.2%—
NYT Connections (extended)—12.9%
CritPt0%—
LiveBench Reasoning65.8%—
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Mistral Medium 3.5 leads

DeepSeek-V3: 32.1 (#219), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Math13731431
OTIS Mock AIME 2024-202537.8%—
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Mistral Medium 3.5 leads

DeepSeek-V3: 37.5 (#155), Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Expert13511432
GPQA Diamond67.6%—
MMLU-Pro72.3%—
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multimodal Not comparable

DeepSeek-V3: —, Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Vision—1223

Multilingual Mistral Medium 3.5 leads

DeepSeek-V3: 48.5 (#143), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Non-English13581404
LMArena Chinese13911442
LMArena French13851448
LMArena German13741451
LMArena Korean13191385
LMArena Russian13731395
LMArena Spanish13581409
LMArena Japanese1333—

Instruction Following Mistral Medium 3.5 leads

DeepSeek-V3: 72.8 (#130), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Instruction Following13451415
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context Mistral Medium 3.5 leads

DeepSeek-V3: 34.0 (#253), Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Longer Query13521415
Fiction.LiveBench50%—

Writing & Preference Mistral Medium 3.5 leads

DeepSeek-V3: 57.4 (#130), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Mistral Medium 3.5
LMArena Text13751421
LMArena Creative Writing13641374
LMArena Multi-Turn13891423
Short-Story Creative Writing77%—
EQ-Bench Creative Writing1472—
WildBench83%—
EQ-Bench 4—993
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Mistral Medium 3.5?

DeepSeek-V3 and Mistral Medium 3.5 score almost the same on the Noometry Index (39.5 vs 40.2), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3 or Mistral Medium 3.5?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is DeepSeek-V3 or Mistral Medium 3.5 better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 36.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-V3 and Mistral Medium 3.5 share?

18 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper