Model comparison

Llama2 70b Steerlm Chat vs Wizardlm 13b

Llama2 70b Steerlm Chat and Wizardlm 13b score almost the same on the Noometry Index (31.8 vs 31.4), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 5 categories and Wizardlm 13b in 2 categories; 4 gaps are clear of the uncertainty.

Side by side

Llama2 70b Steerlm Chat and Wizardlm 13b specifications
Llama2 70b Steerlm ChatWizardlm 13b
ProviderNVIDIAMicrosoft
Noometry Index31.831.4
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked910

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama2 70b Steerlm Chat: 29.9 (#300), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Coding10251035

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Hard Prompts10471018

Math Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.3 (#226), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Math10721017

Multilingual Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 28.8 (#270), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Non-English10631034
LMArena Chinese—1023

Instruction Following Too close to call

Llama2 70b Steerlm Chat: 54.2 (#279), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Instruction Following10601048

Long Context Wizardlm 13b leads

Llama2 70b Steerlm Chat: 30.4 (#288), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Longer Query9981054

Writing & Preference Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.6 (#283), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatWizardlm 13b
LMArena Text10981077
LMArena Creative Writing10911091
LMArena Multi-Turn10581047

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Wizardlm 13b?

Llama2 70b Steerlm Chat and Wizardlm 13b score almost the same on the Noometry Index (31.8 vs 31.4), so choose on price, context window or the category you care about most.

Is Llama2 70b Steerlm Chat or Wizardlm 13b better for coding?

They score almost the same on coding (29.9 vs 30.1); test both on your own repository before choosing.

How many benchmarks do Llama2 70b Steerlm Chat and Wizardlm 13b share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper