Model comparison

Olmo 3.1 32b Instruct vs Wizardlm 70b

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 33.0 on the Noometry Index.

Last verified . 12 shared benchmarks.

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 7 categories and Wizardlm 70b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 3.1 32b Instruct leads 50.2 to 34.8.

Side by side

Olmo 3.1 32b Instruct and Wizardlm 70b specifications
Olmo 3.1 32b InstructWizardlm 70b
ProviderAllen Institute for AI (Ai2)Microsoft
Noometry Index39.433.0
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1612

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Coding13471081

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Hard Prompts13221079

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Math13051116

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Expert1308—

Multilingual Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 42.6 (#191), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Non-English12751078
LMArena Chinese13041052
LMArena German12821083
LMArena Russian12681155
LMArena French1328—
LMArena Korean1206—
LMArena Spanish1336—

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Instruction Following12991093

Long Context Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.9 (#166), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Longer Query13121097

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructWizardlm 70b
LMArena Text13111120
LMArena Creative Writing12641149
LMArena Multi-Turn13091108

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Wizardlm 70b?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 33.0 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Wizardlm 70b better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 31.4 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Wizardlm 70b share?

12 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper