Model comparison

Llama 3.1-405B vs Olmo 7b Instruct

Llama 3.1-405B and Olmo 7b Instruct score almost the same on the Noometry Index (30.7 vs 30.3), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1-405B scores higher in 4 categories and Olmo 7b Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.1-405B leads 65.9 to 49.0.

Side by side

Llama 3.1-405B and Olmo 7b Instruct specifications
Llama 3.1-405BOlmo 7b Instruct
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index30.730.3
Released2024-07-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Llama 3.1-405B: 33.1 (#262), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Coding12911016
WeirdML21.4%—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Olmo 7b Instruct leads

Llama 3.1-405B: 16.8 (#300), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Hard Prompts1269993
SimpleBench23%—
Kagi LLM Benchmark45%—
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Olmo 7b Instruct leads

Llama 3.1-405B: 18.4 (#290), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Math12811018
OTIS Mock AIME 2024-20259.7%—
Omni-MATH24.9%—
MATH Level 549.8%—

Knowledge Not comparable

Llama 3.1-405B: 30.4 (#227), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
GPQA Diamond50.9%—
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
LMArena Expert1243—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Llama 3.1-405B leads

Llama 3.1-405B: 40.7 (#214), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Non-English1248977
LMArena Chinese12421014
LMArena Russian1265947
LMArena French1279—
LMArena German1252—
LMArena Japanese1208—
LMArena Korean1184—
LMArena Spanish1260—

Instruction Following Llama 3.1-405B leads

Llama 3.1-405B: 65.9 (#214), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Instruction Following1259978
IFEval81.1%—

Long Context Not comparable

Llama 3.1-405B: 38.4 (#197), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Longer Query1266—

Writing & Preference Llama 3.1-405B leads

Llama 3.1-405B: 38.9 (#251), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BOlmo 7b Instruct
LMArena Text12841032
LMArena Creative Writing1262990
LMArena Multi-Turn12971007
EQ-Bench Creative Writing870—
WildBench78.3%—

Frequently asked questions

Is Llama 3.1-405B better than Olmo 7b Instruct?

Llama 3.1-405B and Olmo 7b Instruct score almost the same on the Noometry Index (30.7 vs 30.3), so choose on price, context window or the category you care about most.

Is Llama 3.1-405B or Olmo 7b Instruct better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 29.6 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper