Model comparison

Claude 2.1 vs Olmo 7b Instruct

Olmo 7b Instruct is the stronger model overall, scoring 30.3 to 25.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Summary

  • The widest gap is in math, where Olmo 7b Instruct leads 30.2 to 10.2.
  • Olmo 7b Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Olmo 7b Instruct specifications
Claude 2.1Olmo 7b Instruct
ProviderAnthropicAllen Institute for AI (Ai2)
Noometry Index25.230.3
Released2023-11-21—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 7b Instruct leads

Claude 2.1: 26.2 (#327), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
WeirdML7.1%—
LMArena Coding—1016

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
LMArena Hard Prompts—993
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Olmo 7b Instruct leads

Claude 2.1: 10.2 (#315), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
OTIS Mock AIME 2024-20251.9%—
LMArena Math—1018

Knowledge Not comparable

Claude 2.1: 15.4 (#292), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
GPQA Diamond33%—
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
LMArena Non-English—977
LMArena Chinese—1014
LMArena Russian—947

Instruction Following Not comparable

Claude 2.1: —, Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
LMArena Instruction Following—978

Writing & Preference Not comparable

Claude 2.1: —, Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkClaude 2.1Olmo 7b Instruct
LMArena Text—1032
LMArena Creative Writing—990
LMArena Multi-Turn—1007

Frequently asked questions

Is Claude 2.1 better than Olmo 7b Instruct?

Olmo 7b Instruct is the stronger model overall, scoring 30.3 to 25.2 on the Noometry Index.

Is Claude 2.1 or Olmo 7b Instruct better for coding?

Olmo 7b Instruct scores higher on coding benchmarks: 29.6 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Olmo 7b Instruct share?

0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper