Model comparison

Claude 3.5 Sonnet vs Olmo 3.1 32b Instruct

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 34.6 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 3 categories and Olmo 3.1 32b Instruct in 5 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Instruct leads 36.3 to 19.2.
  • Olmo 3.1 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Olmo 3.1 32b Instruct specifications
Claude 3.5 SonnetOlmo 3.1 32b Instruct
ProviderAnthropicAllen Institute for AI (Ai2)
Noometry Index34.639.4
Released2024-06-20—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6016

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Coding13421347
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Olmo 3.1 32b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Olmo 3.1 32b Instruct leads

Claude 3.5 Sonnet: 23.1 (#183), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Hard Prompts13051322
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Olmo 3.1 32b Instruct leads

Claude 3.5 Sonnet: 19.2 (#288), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Math13071305
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Olmo 3.1 32b Instruct leads

Claude 3.5 Sonnet: 28.6 (#245), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Expert12651308
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Olmo 3.1 32b Instruct: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Too close to call

Claude 3.5 Sonnet: 43.2 (#185), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Non-English12831275
LMArena Chinese12721304
LMArena French13051328
LMArena German12971282
LMArena Korean12001206
LMArena Russian13061268
LMArena Spanish12901336
LMArena Japanese1234—

Instruction Following Too close to call

Claude 3.5 Sonnet: 68.8 (#182), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Instruction Following12971299
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Longer Query13111312

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetOlmo 3.1 32b Instruct
LMArena Text12981311
LMArena Creative Writing12921264
LMArena Multi-Turn13261309
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Olmo 3.1 32b Instruct?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Olmo 3.1 32b Instruct better for coding?

They score almost the same on coding (39.0 vs 39.5); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper