Model comparison

DeepSeek LLM 67B vs Wizardlm 13b

Wizardlm 13b is the stronger model overall, scoring 31.4 to 24.9 on the Noometry Index.

Last verified . 10 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 5 categories and Wizardlm 13b in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Wizardlm 13b leads 30.2 to 8.7.

Side by side

DeepSeek LLM 67B and Wizardlm 13b specifications
DeepSeek LLM 67BWizardlm 13b
ProviderDeepSeekMicrosoft
Noometry Index24.931.4
Released2023-11-29—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1510

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek LLM 67B leads

DeepSeek LLM 67B: 31.9 (#278), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Coding10961035

Reasoning Wizardlm 13b leads

DeepSeek LLM 67B: 16.5 (#304), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Hard Prompts10701018
Chess Puzzles0%—
Epoch Capabilities Index110.5—

Math Wizardlm 13b leads

DeepSeek LLM 67B: 8.7 (#324), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Math11081017
OTIS Mock AIME 2024-20250.8%—
MATH Level 56.4%—

Knowledge Not comparable

DeepSeek LLM 67B: 7.0 (#313), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
GPQA Diamond24.6%—

Multilingual DeepSeek LLM 67B leads

DeepSeek LLM 67B: 29.4 (#267), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Non-English10731034
LMArena Chinese11321023

Instruction Following DeepSeek LLM 67B leads

DeepSeek LLM 67B: 55.4 (#277), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Instruction Following10791048

Long Context DeepSeek LLM 67B leads

DeepSeek LLM 67B: 33.1 (#265), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Longer Query10921054

Writing & Preference DeepSeek LLM 67B leads

DeepSeek LLM 67B: 31.6 (#282), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BWizardlm 13b
LMArena Text11051077
LMArena Creative Writing10671091
LMArena Multi-Turn10821047

Frequently asked questions

Is DeepSeek LLM 67B better than Wizardlm 13b?

Wizardlm 13b is the stronger model overall, scoring 31.4 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or Wizardlm 13b better for coding?

DeepSeek LLM 67B scores higher on coding benchmarks: 31.9 versus 30.1 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Wizardlm 13b share?

10 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper