Model comparison

Qwen1.5 4b Chat vs Wizardlm 70b

Wizardlm 70b is the stronger model overall, scoring 33.0 to 28.8 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Qwen1.5 4b Chat scores higher in 0 categories and Wizardlm 70b in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Wizardlm 70b leads 34.8 to 23.8.

Side by side

Qwen1.5 4b Chat and Wizardlm 70b specifications
Qwen1.5 4b ChatWizardlm 70b
ProviderAlibaba (Qwen)Microsoft
Noometry Index28.833.0
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1312

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 70b leads

Qwen1.5 4b Chat: 29.1 (#308), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Coding9991081

Reasoning Wizardlm 70b leads

Qwen1.5 4b Chat: 18.5 (#279), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Hard Prompts9761079

Math Wizardlm 70b leads

Qwen1.5 4b Chat: 30.4 (#234), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Math10261116

Knowledge Not comparable

Qwen1.5 4b Chat: 26.7 (#255), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Expert980—

Multilingual Wizardlm 70b leads

Qwen1.5 4b Chat: 24.1 (#290), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Non-English9791078
LMArena Chinese10241052
LMArena German9021083
LMArena Russian9521155

Instruction Following Wizardlm 70b leads

Qwen1.5 4b Chat: 49.0 (#300), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Instruction Following9781093

Long Context Wizardlm 70b leads

Qwen1.5 4b Chat: 30.1 (#290), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Longer Query9881097

Writing & Preference Wizardlm 70b leads

Qwen1.5 4b Chat: 23.8 (#309), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkQwen1.5 4b ChatWizardlm 70b
LMArena Text9971120
LMArena Creative Writing9691149
LMArena Multi-Turn9771108

Frequently asked questions

Is Qwen1.5 4b Chat better than Wizardlm 70b?

Wizardlm 70b is the stronger model overall, scoring 33.0 to 28.8 on the Noometry Index.

Is Qwen1.5 4b Chat or Wizardlm 70b better for coding?

Wizardlm 70b scores higher on coding benchmarks: 31.4 versus 29.1 in the Noometry coding category.

How many benchmarks do Qwen1.5 4b Chat and Wizardlm 70b share?

12 benchmarks have published results for both models. Qwen1.5 4b Chat has 13 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper