Model comparison

Qwen1.5 4b Chat vs Wizardlm 13b

Wizardlm 13b is the stronger model overall, scoring 31.4 to 28.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Qwen1.5 4b Chat scores higher in 1 category and Wizardlm 13b in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Wizardlm 13b leads 30.5 to 23.8.

Side by side

Qwen1.5 4b Chat and Wizardlm 13b specifications
Qwen1.5 4b ChatWizardlm 13b
ProviderAlibaba (Qwen)Microsoft
Noometry Index28.831.4
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 13b leads

Qwen1.5 4b Chat: 29.1 (#308), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Coding9991035

Reasoning Too close to call

Qwen1.5 4b Chat: 18.5 (#279), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Hard Prompts9761018

Math Too close to call

Qwen1.5 4b Chat: 30.4 (#234), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Math10261017

Knowledge Not comparable

Qwen1.5 4b Chat: 26.7 (#255), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Expert980—

Multilingual Wizardlm 13b leads

Qwen1.5 4b Chat: 24.1 (#290), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Non-English9791034
LMArena Chinese10241023
LMArena German902—
LMArena Russian952—

Instruction Following Wizardlm 13b leads

Qwen1.5 4b Chat: 49.0 (#300), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Instruction Following9781048

Long Context Wizardlm 13b leads

Qwen1.5 4b Chat: 30.1 (#290), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Longer Query9881054

Writing & Preference Wizardlm 13b leads

Qwen1.5 4b Chat: 23.8 (#309), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 13b
LMArena Text9971077
LMArena Creative Writing9691091
LMArena Multi-Turn9771047

Frequently asked questions

Is Qwen1.5 4b Chat better than Wizardlm 13b?

Wizardlm 13b is the stronger model overall, scoring 31.4 to 28.8 on the Noometry Index.

Is Qwen1.5 4b Chat or Wizardlm 13b better for coding?

Wizardlm 13b scores higher on coding benchmarks: 30.1 versus 29.1 in the Noometry coding category.

How many benchmarks do Qwen1.5 4b Chat and Wizardlm 13b share?

10 benchmarks have published results for both models. Qwen1.5 4b Chat has 13 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper