Model comparison

Olmo 3.1 32b Think vs Phi-4 Mini

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 30.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in knowledge, where Olmo 3.1 32b Think leads 35.7 to 25.3.

Side by side

Olmo 3.1 32b Think and Phi-4 Mini specifications
Olmo 3.1 32b ThinkPhi-4 Mini
ProviderAllen Institute for AI (Ai2)Microsoft
Noometry Index37.930.9
Released—2024-12-11
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.075
Output $ / M tokens—$0.30
Results tracked153

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
SciCode—10.8%
LMArena Coding1291—

Reasoning Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 25.2 (#150), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
CritPt—0%
LMArena Hard Prompts1272—

Math Not comparable

Olmo 3.1 32b Think: 36.3 (#168), Phi-4 Mini: —

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
LMArena Math1305—

Knowledge Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 35.7 (#181), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
Vectara Hallucination Rate—23.5%
LMArena Expert1295—

Multilingual Not comparable

Olmo 3.1 32b Think: 38.1 (#231), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
LMArena Non-English1209—
LMArena Chinese1242—
LMArena French1260—
LMArena German1262—
LMArena Russian1193—
LMArena Spanish1289—

Instruction Following Not comparable

Olmo 3.1 32b Think: 65.6 (#218), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
LMArena Instruction Following1247—

Long Context Not comparable

Olmo 3.1 32b Think: 38.6 (#195), Phi-4 Mini: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
LMArena Longer Query1272—

Writing & Preference Not comparable

Olmo 3.1 32b Think: 46.2 (#220), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkPhi-4 Mini
LMArena Text1272—
LMArena Creative Writing1226—
LMArena Multi-Turn1252—

Frequently asked questions

Is Olmo 3.1 32b Think better than Phi-4 Mini?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 30.9 on the Noometry Index.

Is Olmo 3.1 32b Think or Phi-4 Mini better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 28.1 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Phi-4 Mini share?

0 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper