Model comparison

Llama 3.1-70B vs Olmo 7b Instruct

Llama 3.1-70B and Olmo 7b Instruct score almost the same on the Noometry Index (29.6 vs 30.3), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1-70B scores higher in 5 categories and Olmo 7b Instruct in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 7b Instruct leads 30.2 to 13.5.

Side by side

Llama 3.1-70B and Olmo 7b Instruct specifications
Llama 3.1-70BOlmo 7b Instruct
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index29.630.3
Released2024-07-23—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.40—
Output $ / M tokens$0.40—
Results tracked3510

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.1-70B: 30.3 (#296), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Coding12601016
WeirdML9%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete54.8%—

Agentic & Tool Use Not comparable

Llama 3.1-70B: 25.1 (#112), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
TheAgentCompany6.9%—
BALROG27.9%—

Reasoning Llama 3.1-70B leads

Llama 3.1-70B: 21.6 (#220), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Hard Prompts1241993
DTBench60%—
LMCA14.8%—
Epoch Capabilities Index125.92—

Math Olmo 7b Instruct leads

Llama 3.1-70B: 13.5 (#304), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Math12521018
OTIS Mock AIME 2024-20253.6%—
Omni-MATH21%—
MATH Level 536.7%—

Knowledge Not comparable

Llama 3.1-70B: 24.2 (#269), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
GPQA Diamond44.2%—
MMLU-Pro65.3%—
GPQA (HELM)42.6%—
LMArena Expert1209—
MMLU80.1%—

Multilingual Llama 3.1-70B leads

Llama 3.1-70B: 38.8 (#225), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Non-English1219977
LMArena Chinese12151014
LMArena Russian1234947
LMArena French1261—
LMArena German1222—
LMArena Japanese1132—
LMArena Korean1140—
LMArena Spanish1253—

Instruction Following Llama 3.1-70B leads

Llama 3.1-70B: 65.3 (#223), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Instruction Following1231978
IFEval82.1%—

Long Context Not comparable

Llama 3.1-70B: 37.6 (#214), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Longer Query1241—

Writing & Preference Llama 3.1-70B leads

Llama 3.1-70B: 35.4 (#267), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkLlama 3.1-70BOlmo 7b Instruct
LMArena Text12611032
LMArena Creative Writing1232990
LMArena Multi-Turn12561007
EQ-Bench Creative Writing784—
WildBench75.8%—

Frequently asked questions

Is Llama 3.1-70B better than Olmo 7b Instruct?

Llama 3.1-70B and Olmo 7b Instruct score almost the same on the Noometry Index (29.6 vs 30.3), so choose on price, context window or the category you care about most.

Is Llama 3.1-70B or Olmo 7b Instruct better for coding?

They score almost the same on coding (30.3 vs 29.6); test both on your own repository before choosing.

How many benchmarks do Llama 3.1-70B and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Llama 3.1-70B has 35 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper