Model comparison

Olmo 3.1 32b Instruct vs Phi-3.5-mini

Olmo 3.1 32b Instruct has enough public results to be ranked (#168); Phi-3.5-mini does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in coding, where Olmo 3.1 32b Instruct leads 39.5 to 34.6.

Side by side

Olmo 3.1 32b Instruct and Phi-3.5-mini specifications
Olmo 3.1 32b InstructPhi-3.5-mini
ProviderAllen Institute for AI (Ai2)Microsoft
Noometry Index39.436.9
Released—2024-08-16
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked165

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Phi-3.5-mini: 34.6 (#232)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
BigCodeBench Instruct—32.8%
LMArena Coding1347—
BigCodeBench Complete—38.5%

Reasoning Not comparable

Olmo 3.1 32b Instruct: 26.4 (#132), Phi-3.5-mini: —

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Hard Prompts1322—
PIQA—81%

Math Not comparable

Olmo 3.1 32b Instruct: 36.3 (#167), Phi-3.5-mini: —

Math benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Math1305—
GSM8K—86.2%

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Phi-3.5-mini: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Expert1308—
BoolQ—78%

Multilingual Not comparable

Olmo 3.1 32b Instruct: 42.6 (#191), Phi-3.5-mini: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Non-English1275—
LMArena Chinese1304—
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Russian1268—
LMArena Spanish1336—

Instruction Following Not comparable

Olmo 3.1 32b Instruct: 68.6 (#187), Phi-3.5-mini: —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Instruction Following1299—

Long Context Not comparable

Olmo 3.1 32b Instruct: 39.9 (#166), Phi-3.5-mini: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Longer Query1312—

Writing & Preference Not comparable

Olmo 3.1 32b Instruct: 50.2 (#185), Phi-3.5-mini: —

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructPhi-3.5-mini
LMArena Text1311—
LMArena Creative Writing1264—
LMArena Multi-Turn1309—

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Phi-3.5-mini?

Olmo 3.1 32b Instruct has enough public results to be ranked (#168); Phi-3.5-mini does not yet, so treat this comparison as directional.

Is Olmo 3.1 32b Instruct or Phi-3.5-mini better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 34.6 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Phi-3.5-mini share?

0 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Phi-3.5-mini has 5.

Related comparisons

Go deeper