Model comparison

DeepSeek V4 Flash vs Mistral Medium

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 36.3 on the Noometry Index.

Last verified . 28 shared benchmarks.

DeepSeek V4 Flash DeepSeek

53.6

Rank #35 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 28 benchmarks with published results for both. DeepSeek V4 Flash scores higher in 8 categories and Mistral Medium in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4 Flash leads 60.3 to 28.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 94.4% for DeepSeek V4 Flash and 32.2% for Mistral Medium.
  • DeepSeek V4 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • DeepSeek V4 Flash accepts more context: 1M tokens versus 262K.

Side by side

DeepSeek V4 Flash and Mistral Medium specifications
DeepSeek V4 FlashMistral Medium
ProviderDeepSeekMistral AI
Noometry Index53.636.3
Released2026-04-242023-12-11
WeightsOpenOpen
Context window1M262K
Max output393K262K
Input $ / M tokens$0.15$1.50
Output $ / M tokens$0.60$7.50
Results tracked4136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Flash leads

DeepSeek V4 Flash: 47.9 (#59), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
FrontierCode18.8%8%
SciCode49.9%40.2%
WeirdML63%43.7%
LMArena Coding14571434
ALE-Bench1,306763.98
LMArena WebDev1582—

Agentic & Tool Use Not comparable

DeepSeek V4 Flash: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning DeepSeek V4 Flash leads

DeepSeek V4 Flash: 53.7 (#30), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
Kagi LLM Benchmark52.2%50%
CritPt16.6%0%
LMArena Hard Prompts14441426
DTBench90.9%75.5%
LMCA41.7%26.1%
ARC-AGI-261.4%—
SimpleBench61.1%—
NYT Connections (extended)89.6%—
ARC-AGI-189%—
Chess Puzzles33%—
Mystery Game Puzzles34%—
Surface Evolver Bench—26.9%
Epoch Capabilities Index154.49—

Math DeepSeek V4 Flash leads

DeepSeek V4 Flash: 60.3 (#37), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
OTIS Mock AIME 2024-202594.4%32.2%
ProofBench56%9%
LMArena Math14271408
FrontierMath (Tiers 1-3)57.5%—
FrontierMath Tier 424.4%—
MathArena Final-Answer Competitions76.5%—
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge DeepSeek V4 Flash leads

DeepSeek V4 Flash: 55.4 (#48), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
GPQA Diamond91%59.5%
LMArena Expert14411408
Humanity's Last Exam—4.5%
SimpleQA Verified33.6%—
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

DeepSeek V4 Flash: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
LMArena Vision—1172

Multilingual Too close to call

DeepSeek V4 Flash: 53.0 (#72), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
LMArena Non-English14201408
LMArena Chinese14681447
LMArena French14391459
LMArena German14181432
LMArena Japanese14061378
LMArena Korean13841380
LMArena Russian14281411
LMArena Spanish14361433

Instruction Following DeepSeek V4 Flash leads

DeepSeek V4 Flash: 74.9 (#81), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
LMArena Instruction Following14211398

Long Context Too close to call

DeepSeek V4 Flash: 43.8 (#85), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
LMArena Longer Query14341406

Writing & Preference DeepSeek V4 Flash leads

DeepSeek V4 Flash: 63.8 (#61), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 FlashMistral Medium
LMArena Text14321424
LMArena Creative Writing14031391
LMArena Multi-Turn14491418
Short-Story Creative Writing—77.3%
EQ-Bench Creative Writing1559—

Frequently asked questions

Is DeepSeek V4 Flash better than Mistral Medium?

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 36.3 on the Noometry Index.

Which is cheaper, DeepSeek V4 Flash or Mistral Medium?

DeepSeek V4 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is DeepSeek V4 Flash or Mistral Medium better for coding?

DeepSeek V4 Flash scores higher on coding benchmarks: 47.9 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

DeepSeek V4 Flash does, with 1M tokens against 262K.

How many benchmarks do DeepSeek V4 Flash and Mistral Medium share?

28 benchmarks have published results for both models. DeepSeek V4 Flash has 41 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper