Model comparison

Olmo 3.1 32b Instruct vs Wizardlm 13b

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 7 categories and Wizardlm 13b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 3.1 32b Instruct leads 50.2 to 30.5.

Side by side

Olmo 3.1 32b Instruct and Wizardlm 13b specifications
Olmo 3.1 32b InstructWizardlm 13b
ProviderAllen Institute for AI (Ai2)Microsoft
Noometry Index39.431.4
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1610

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Coding13471035

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Hard Prompts13221018

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Math13051017

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Expert1308—

Multilingual Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 42.6 (#191), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Non-English12751034
LMArena Chinese13041023
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Russian1268—
LMArena Spanish1336—

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Instruction Following12991048

Long Context Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.9 (#166), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Longer Query13121054

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 13b
LMArena Text13111077
LMArena Creative Writing12641091
LMArena Multi-Turn13091047

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Wizardlm 13b?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 31.4 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Wizardlm 13b better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 30.1 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Wizardlm 13b share?

10 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper