Model comparison

DeepSeek LLM 67B vs Mistral Medium 3.1

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 24.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 31.6.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and Mistral Medium 3.1 specifications
DeepSeek LLM 67BMistral Medium 3.1
ProviderDeepSeekMistral AI
Noometry Index24.931.9
Released2023-11-29—
WeightsOpenProprietary
Context window—131K
Max output—105K
Input $ / M tokens—$0.40
Output $ / M tokens—$2
Results tracked153

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek LLM 67B: 31.9 (#278), Mistral Medium 3.1: —

Coding benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
LMArena Coding1096—

Reasoning DeepSeek LLM 67B leads

DeepSeek LLM 67B: 16.5 (#304), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
NYT Connections (extended)—6.5%
Chess Puzzles0%—
Thematic Generalization—20.3%
LMArena Hard Prompts1070—
Epoch Capabilities Index110.5—

Math Not comparable

DeepSeek LLM 67B: 8.7 (#324), Mistral Medium 3.1: —

Math benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
OTIS Mock AIME 2024-20250.8%—
LMArena Math1108—
MATH Level 56.4%—

Knowledge Not comparable

DeepSeek LLM 67B: 7.0 (#313), Mistral Medium 3.1: —

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
GPQA Diamond24.6%—

Multilingual Not comparable

DeepSeek LLM 67B: 29.4 (#267), Mistral Medium 3.1: —

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
LMArena Non-English1073—
LMArena Chinese1132—

Instruction Following Not comparable

DeepSeek LLM 67B: 55.4 (#277), Mistral Medium 3.1: —

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
LMArena Instruction Following1079—

Long Context Not comparable

DeepSeek LLM 67B: 33.1 (#265), Mistral Medium 3.1: —

Long Context benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
LMArena Longer Query1092—

Writing & Preference Mistral Medium 3.1 leads

DeepSeek LLM 67B: 31.6 (#282), Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BMistral Medium 3.1
LMArena Text1105—
LMArena Creative Writing1067—
EQ-Bench Creative Writing—1476
LMArena Multi-Turn1082—

Frequently asked questions

Is DeepSeek LLM 67B better than Mistral Medium 3.1?

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 24.9 on the Noometry Index.

How many benchmarks do DeepSeek LLM 67B and Mistral Medium 3.1 share?

0 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper