Model comparison

Olmo 3.1 32b Think vs Wizardlm 13b

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Olmo 3.1 32b Think scores higher in 7 categories and Wizardlm 13b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 3.1 32b Think leads 46.2 to 30.5.

Side by side

Olmo 3.1 32b Think and Wizardlm 13b specifications
Olmo 3.1 32b ThinkWizardlm 13b
ProviderAllen Institute for AI (Ai2)Microsoft
Noometry Index37.931.4
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1510

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Coding12911035

Reasoning Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 25.2 (#150), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Hard Prompts12721018

Math Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 36.3 (#168), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Math13051017

Knowledge Not comparable

Olmo 3.1 32b Think: 35.7 (#181), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Expert1295—

Multilingual Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 38.1 (#231), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Non-English12091034
LMArena Chinese12421023
LMArena French1260—
LMArena German1262—
LMArena Russian1193—
LMArena Spanish1289—

Instruction Following Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 65.6 (#218), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Instruction Following12471048

Long Context Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 38.6 (#195), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Longer Query12721054

Writing & Preference Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 46.2 (#220), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkWizardlm 13b
LMArena Text12721077
LMArena Creative Writing12261091
LMArena Multi-Turn12521047

Frequently asked questions

Is Olmo 3.1 32b Think better than Wizardlm 13b?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 31.4 on the Noometry Index.

Is Olmo 3.1 32b Think or Wizardlm 13b better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 30.1 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Wizardlm 13b share?

10 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper