Model comparison

DeepSeek Coder 33B vs Mixtral 8x22B

Mixtral 8x22B has enough public results to be ranked (#333); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Last verified . 6 shared benchmarks.

DeepSeek Coder 33B DeepSeek

38.9

Unranked Sparse

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 6 benchmarks with published results for both. DeepSeek Coder 33B scores higher in 1 category and Mixtral 8x22B in 0 categories; one gap is clear of the uncertainty.
  • The widest gap is in coding, where DeepSeek Coder 33B leads 38.0 to 24.2.

Side by side

DeepSeek Coder 33B and Mixtral 8x22B specifications
DeepSeek Coder 33BMixtral 8x22B
ProviderDeepSeekMistral AI
Noometry Index38.927.1
Released2023-11-022024-04-17
WeightsOpenOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek Coder 33B leads

DeepSeek Coder 33B: 38.0 (#184), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
BigCodeBench Instruct42%40.6%
BigCodeBench Complete51.1%50.2%
HumanEval+75%72%
MBPP+70.1%64.3%
WeirdML—3.2%
LMArena Coding—1166

Agentic & Tool Use Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
Cybench—7.5%

Reasoning Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
Epoch Capabilities Index96.32122.03
LMArena Hard Prompts—1150
DTBench—55.1%
ForecastBench—56.3
WinoGrande62%—

Math Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
Omni-MATH—16.3%
LMArena Math—1184
MATH Level 5—24.2%
GSM8K35.4%—

Knowledge Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
MMLU39.4%77.8%
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
ARC (AI2) Challenge42.2%—

Multilingual Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
LMArena Non-English—1128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
IFEval—72.4%
LMArena Instruction Following—1147

Long Context Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
LMArena Longer Query—1144

Writing & Preference Not comparable

DeepSeek Coder 33B: —, Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkDeepSeek Coder 33BMixtral 8x22B
LMArena Text—1162
LMArena Creative Writing—1141
WildBench—71.1%
LMArena Multi-Turn—1130

Frequently asked questions

Is DeepSeek Coder 33B better than Mixtral 8x22B?

Mixtral 8x22B has enough public results to be ranked (#333); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Is DeepSeek Coder 33B or Mixtral 8x22B better for coding?

DeepSeek Coder 33B scores higher on coding benchmarks: 38.0 versus 24.2 in the Noometry coding category.

How many benchmarks do DeepSeek Coder 33B and Mixtral 8x22B share?

6 benchmarks have published results for both models. DeepSeek Coder 33B has 9 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper