Model comparison

Olmo 7b Instruct vs Phi 3 Mini 128k Instruct

Olmo 7b Instruct and Phi 3 Mini 128k Instruct score almost the same on the Noometry Index (30.3 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Summary

  • They share 10 benchmarks with published results for both. Olmo 7b Instruct scores higher in 1 category and Phi 3 Mini 128k Instruct in 5 categories; 4 gaps are clear of the uncertainty.

Side by side

Olmo 7b Instruct and Phi 3 Mini 128k Instruct specifications
Olmo 7b InstructPhi 3 Mini 128k Instruct
ProviderAllen Institute for AI (Ai2)Microsoft
Noometry Index30.329.7
Released—2024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Olmo 7b Instruct: 29.6 (#303), Phi 3 Mini 128k Instruct: 28.8 (#312)

Coding benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Coding10161039
BigCodeBench Instruct—29.6%
BigCodeBench Complete—40.6%

Reasoning Too close to call

Olmo 7b Instruct: 18.8 (#274), Phi 3 Mini 128k Instruct: 19.6 (#256)

Reasoning benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Hard Prompts9931028

Math Phi 3 Mini 128k Instruct leads

Olmo 7b Instruct: 30.2 (#237), Phi 3 Mini 128k Instruct: 31.6 (#222)

Math benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Math10181089

Knowledge Not comparable

Olmo 7b Instruct: —, Phi 3 Mini 128k Instruct: 26.8 (#254)

Knowledge benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Expert—984

Multilingual Phi 3 Mini 128k Instruct leads

Olmo 7b Instruct: 24.0 (#291), Phi 3 Mini 128k Instruct: 25.2 (#285)

Multilingual benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Non-English9771000
LMArena Chinese10141016
LMArena Russian9471004
LMArena French—1039
LMArena German—1006
LMArena Japanese—899
LMArena Korean—856
LMArena Spanish—1059

Instruction Following Phi 3 Mini 128k Instruct leads

Olmo 7b Instruct: 49.0 (#301), Phi 3 Mini 128k Instruct: 51.8 (#294)

Instruction Following benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Instruction Following9781023

Long Context Not comparable

Olmo 7b Instruct: —, Phi 3 Mini 128k Instruct: 30.4 (#289)

Long Context benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Longer Query—996

Writing & Preference Phi 3 Mini 128k Instruct leads

Olmo 7b Instruct: 25.8 (#303), Phi 3 Mini 128k Instruct: 27.1 (#301)

Writing & Preference benchmarks
BenchmarkOlmo 7b InstructPhi 3 Mini 128k Instruct
LMArena Text10321050
LMArena Creative Writing9901024
LMArena Multi-Turn1007989

Frequently asked questions

Is Olmo 7b Instruct better than Phi 3 Mini 128k Instruct?

Olmo 7b Instruct and Phi 3 Mini 128k Instruct score almost the same on the Noometry Index (30.3 vs 29.7), so choose on price, context window or the category you care about most.

Is Olmo 7b Instruct or Phi 3 Mini 128k Instruct better for coding?

They score almost the same on coding (29.6 vs 28.8); test both on your own repository before choosing.

How many benchmarks do Olmo 7b Instruct and Phi 3 Mini 128k Instruct share?

10 benchmarks have published results for both models. Olmo 7b Instruct has 10 scored results on Noometry and Phi 3 Mini 128k Instruct has 19.

Related comparisons

Go deeper