Model comparison

Claude 3.5 Sonnet vs Olmo 7b Instruct

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 5 categories and Olmo 7b Instruct in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude 3.5 Sonnet leads 52.9 to 25.8.
  • Olmo 7b Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Olmo 7b Instruct specifications
Claude 3.5 SonnetOlmo 7b Instruct
ProviderAnthropicAllen Institute for AI (Ai2)
Noometry Index34.630.3
Released2024-06-20—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6010

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Coding13421016
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Olmo 7b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Hard Prompts1305993
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Olmo 7b Instruct leads

Claude 3.5 Sonnet: 19.2 (#288), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Math13071018
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Not comparable

Claude 3.5 Sonnet: 28.6 (#245), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
LMArena Expert1265—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Olmo 7b Instruct: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 43.2 (#185), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Non-English1283977
LMArena Chinese12721014
LMArena Russian1306947
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Spanish1290—

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Instruction Following1297978
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Not comparable

Claude 3.5 Sonnet: 39.9 (#167), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Longer Query1311—

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetOlmo 7b Instruct
LMArena Text12981032
LMArena Creative Writing1292990
LMArena Multi-Turn13261007
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Olmo 7b Instruct?

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 30.3 on the Noometry Index.

Is Claude 3.5 Sonnet or Olmo 7b Instruct better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 29.6 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper