Model comparison

DeepSeek Coder 33B vs Wizardlm 13b

Wizardlm 13b has enough public results to be ranked (#274); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

DeepSeek Coder 33B DeepSeek

38.9

Unranked Sparse

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • The widest gap is in coding, where DeepSeek Coder 33B leads 38.0 to 30.1.

Side by side

DeepSeek Coder 33B and Wizardlm 13b specifications
DeepSeek Coder 33BWizardlm 13b
ProviderDeepSeekMicrosoft
Noometry Index38.931.4
Released2023-11-02—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked910

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek Coder 33B leads

DeepSeek Coder 33B: 38.0 (#184), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
BigCodeBench Instruct42%—
LMArena Coding—1035
BigCodeBench Complete51.1%—
HumanEval+75%—
MBPP+70.1%—

Reasoning Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
LMArena Hard Prompts—1018
Epoch Capabilities Index96.32—
WinoGrande62%—

Math Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
LMArena Math—1017
GSM8K35.4%—

Knowledge Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: —

Knowledge benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
ARC (AI2) Challenge42.2%—
MMLU39.4%—

Multilingual Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
LMArena Non-English—1034
LMArena Chinese—1023

Instruction Following Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
LMArena Instruction Following—1048

Long Context Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
LMArena Longer Query—1054

Writing & Preference Not comparable

DeepSeek Coder 33B: —, Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkDeepSeek Coder 33BWizardlm 13b
LMArena Text—1077
LMArena Creative Writing—1091
LMArena Multi-Turn—1047

Frequently asked questions

Is DeepSeek Coder 33B better than Wizardlm 13b?

Wizardlm 13b has enough public results to be ranked (#274); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Is DeepSeek Coder 33B or Wizardlm 13b better for coding?

DeepSeek Coder 33B scores higher on coding benchmarks: 38.0 versus 30.1 in the Noometry coding category.

How many benchmarks do DeepSeek Coder 33B and Wizardlm 13b share?

0 benchmarks have published results for both models. DeepSeek Coder 33B has 9 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper