Model comparison

Llama 2-13B vs Olmo 7b Instruct

Llama 2-13B and Olmo 7b Instruct score almost the same on the Noometry Index (29.6 vs 30.3), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Llama 2-13B Meta

29.6

Rank #309 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 2-13B scores higher in 5 categories and Olmo 7b Instruct in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 7b Instruct leads 18.8 to 12.8.

Side by side

Llama 2-13B and Olmo 7b Instruct specifications
Llama 2-13BOlmo 7b Instruct
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index29.630.3
Released2023-07-18—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 2-13B leads

Llama 2-13B: 30.9 (#291), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Coding10621016

Reasoning Olmo 7b Instruct leads

Llama 2-13B: 12.8 (#337), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Hard Prompts1051993
Chess Puzzles0%—
DTBench42.2%—
BIG-Bench Hard58.2%—
Epoch Capabilities Index106.17—
HellaSwag80.7%—
LAMBADA76.5%—
PIQA80.8%—
WinoGrande72.8%—

Math Too close to call

Llama 2-13B: 31.1 (#229), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Math10651018
GSM8K36.9%—

Knowledge Not comparable

Llama 2-13B: 28.1 (#249), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Expert1030—
ARC (AI2) Challenge60.3%—
BoolQ82.4%—
MMLU55.6%—
OpenBookQA57%—
TriviaQA79.6%—

Multimodal Not comparable

Llama 2-13B: —, Olmo 7b Instruct: —

Multimodal benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
ScienceQA55.8%—

Multilingual Llama 2-13B leads

Llama 2-13B: 26.5 (#279), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Non-English1024977
LMArena Chinese10011014
LMArena Russian1055947
LMArena French1044—
LMArena German1009—
LMArena Japanese894—
LMArena Korean953—
LMArena Spanish1087—

Instruction Following Llama 2-13B leads

Llama 2-13B: 53.3 (#287), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Instruction Following1045978

Long Context Not comparable

Llama 2-13B: 32.3 (#269), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Longer Query1064—

Writing & Preference Llama 2-13B leads

Llama 2-13B: 29.8 (#289), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkLlama 2-13BOlmo 7b Instruct
LMArena Text10841032
LMArena Creative Writing1047990
LMArena Multi-Turn10501007

Frequently asked questions

Is Llama 2-13B better than Olmo 7b Instruct?

Llama 2-13B and Olmo 7b Instruct score almost the same on the Noometry Index (29.6 vs 30.3), so choose on price, context window or the category you care about most.

Is Llama 2-13B or Olmo 7b Instruct better for coding?

Llama 2-13B scores higher on coding benchmarks: 30.9 versus 29.6 in the Noometry coding category.

How many benchmarks do Llama 2-13B and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Llama 2-13B has 32 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper