Model comparison

DeepSeek-V2.5 (Sep 2024) vs Mistral Large 3

Mistral Large 3 is the stronger model overall, scoring 39.1 to 37.6 on the Noometry Index.

Last verified . 17 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 1 category and Mistral Large 3 in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V2.5 (Sep 2024) leads 25.6 to 15.2.

Side by side

DeepSeek-V2.5 (Sep 2024) and Mistral Large 3 specifications
DeepSeek-V2.5 (Sep 2024)Mistral Large 3
ProviderDeepSeekMistral AI
Noometry Index37.639.1
Released2024-09-062025-12-02
WeightsOpenOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Coding13091448
Aider Polyglot17.8%—
LMArena WebDev—1230
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Hard Prompts12891429
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
Thematic Generalization—23%

Math Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Math12881414

Knowledge Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Expert12661421
Vectara Hallucination Rate—14.5%

Multimodal Not comparable

DeepSeek-V2.5 (Sep 2024): —, Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Vision—1221

Multilingual Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Non-English12731413
LMArena Chinese13181447
LMArena French12891455
LMArena German12581437
LMArena Japanese12281394
LMArena Korean12091384
LMArena Russian12891411
LMArena Spanish12481440

Instruction Following Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Instruction Following12801403

Long Context Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Longer Query13011413

Writing & Preference Mistral Large 3 leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Mistral Large 3
LMArena Text12941428
LMArena Creative Writing12851386
LMArena Multi-Turn12971429
EQ-Bench Creative Writing—1412

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Mistral Large 3?

Mistral Large 3 is the stronger model overall, scoring 39.1 to 37.6 on the Noometry Index.

Is DeepSeek-V2.5 (Sep 2024) or Mistral Large 3 better for coding?

Mistral Large 3 scores higher on coding benchmarks: 34.4 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Mistral Large 3 share?

17 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper