Model comparison

Codestral vs DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 30.6 on the Noometry Index.

Last verified . 1 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

Summary

  • They share 1 benchmark with published results for both. Codestral scores higher in 0 categories and DeepSeek V4.1 Flash in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4.1 Flash leads 50.2 to 19.8.
  • DeepSeek V4.1 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.30 / $0.90 for Codestral.
  • DeepSeek V4.1 Flash accepts more context: 1M tokens versus 256K.
  • DeepSeek V4.1 Flash has downloadable open weights; the other is API-only.

Side by side

Codestral and DeepSeek V4.1 Flash specifications
CodestralDeepSeek V4.1 Flash
ProviderMistral AIDeepSeek
Noometry Index30.652.8
Released2024-05-292026-09-09
WeightsProprietaryOpen
Context window256K1M
Max output8K393K
Input $ / M tokens$0.30$0.15
Output $ / M tokens$0.90$0.60
Results tracked737

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4.1 Flash leads

Codestral: 27.3 (#321), DeepSeek V4.1 Flash: 52.9 (#32)

Coding benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
ALE-Bench137.781,092
Aider Polyglot11.1%—
LMArena WebDev—1619
SciCode—51.9%
BigCodeBench Instruct41.8%—
LMArena Coding—1506
BigCodeBench Complete52.5%—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, DeepSeek V4.1 Flash: 31.2 (#69)

Agentic & Tool Use benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
APEX-Agents—39.5%
GDP.pdf—19.8%

Reasoning DeepSeek V4.1 Flash leads

Codestral: 19.8 (#251), DeepSeek V4.1 Flash: 50.2 (#36)

Reasoning benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
Kagi LLM Benchmark32.5%—
NYT Connections (extended)—89.6%
CritPt—14.3%
LMArena Hard Prompts—1483
Mystery Game Puzzles—43%
DTBench—89.9%
LMCA—47%
Surface Evolver Bench—46.3%
Epoch Capabilities Index—154.9

Math Not comparable

Codestral: —, DeepSeek V4.1 Flash: 66.7 (#25)

Math benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
FrontierMath (Tiers 1-3)—67.4%
FrontierMath Tier 4—26.8%
OTIS Mock AIME 2024-2025—98.3%
ProofBench—54%
LMArena Math—1477

Knowledge Not comparable

Codestral: —, DeepSeek V4.1 Flash: 57.9 (#38)

Knowledge benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
GPQA Diamond—89.8%
LMArena Expert—1506

Multimodal Not comparable

Codestral: —, DeepSeek V4.1 Flash: 39.1 (#61)

Multimodal benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
LMArena Vision—1277
Furniture Assembly—34.2%

Multilingual Not comparable

Codestral: —, DeepSeek V4.1 Flash: 55.0 (#35)

Multilingual benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
LMArena Non-English—1448
LMArena Chinese—1497
LMArena French—1452
LMArena German—1484
LMArena Japanese—1412
LMArena Korean—1452
LMArena Russian—1471
LMArena Spanish—1459

Instruction Following Not comparable

Codestral: —, DeepSeek V4.1 Flash: 77.3 (#26)

Instruction Following benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
LMArena Instruction Following—1474

Long Context Not comparable

Codestral: —, DeepSeek V4.1 Flash: 45.2 (#47)

Long Context benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
LMArena Longer Query—1475

Writing & Preference Not comparable

Codestral: —, DeepSeek V4.1 Flash: 65.4 (#48)

Writing & Preference benchmarks
BenchmarkCodestralDeepSeek V4.1 Flash
LMArena Text—1462
LMArena Creative Writing—1435
EQ-Bench Creative Writing—1540
LMArena Multi-Turn—1457

Frequently asked questions

Is Codestral better than DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 30.6 on the Noometry Index.

Which is cheaper, Codestral or DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Codestral lists at $0.30 and $0.90.

Is Codestral or DeepSeek V4.1 Flash better for coding?

DeepSeek V4.1 Flash scores higher on coding benchmarks: 52.9 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

DeepSeek V4.1 Flash does, with 1M tokens against 256K.

How many benchmarks do Codestral and DeepSeek V4.1 Flash share?

1 benchmark has published results for both models. Codestral has 7 scored results on Noometry and DeepSeek V4.1 Flash has 37.

Related comparisons

Go deeper