Model comparison

Claude 3.5 Sonnet vs o1

o1 is the stronger model overall, scoring 40.9 to 34.6 on the Noometry Index.

Last verified . 45 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 45 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 1 category and o1 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1 leads 36.1 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 73.3% for o1.

Side by side

Claude 3.5 Sonnet and o1 specifications
Claude 3.5 Sonneto1
ProviderAnthropicOpenAI
Noometry Index34.640.9
Released2024-06-202024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked6052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Claude 3.5 Sonnet: 39.0 (#165), o1: 46.1 (#70)

Coding benchmarks
BenchmarkClaude 3.5 Sonneto1
Aider Polyglot51.6%61.7%
WeirdML40%47.6%
LiveBench Coding67.1%69.7%
LMArena Coding13421367
CadEval48%56%
HumanEval+81.7%89%
MBPP+74.3%80.2%
GSO4.6%—
BigCodeBench Instruct46.8%—
BigCodeBench Complete58.6%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 Sonneto1
Cybench17.5%10%
METR Time Horizons45.2%51.1%
TheAgentCompany24%—
BALROG32.6%—

Reasoning o1 leads

Claude 3.5 Sonnet: 23.1 (#183), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkClaude 3.5 Sonneto1
SimpleBench41.4%41.7%
EnigmaEval0.9%5.7%
LiveBench Reasoning56.7%91.6%
LMArena Hard Prompts13051371
DTBench67.8%74.7%
LiveBench Data Analysis55%65.5%
Epoch Capabilities Index133.55141.91
LiveBench59%75.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
LMCA—22.3%
ForecastBench60.7—

Math o1 leads

Claude 3.5 Sonnet: 19.2 (#288), o1: 36.1 (#175)

Math benchmarks
BenchmarkClaude 3.5 Sonneto1
OTIS Mock AIME 2024-20258.5%73.3%
LiveBench Math52.3%80.3%
LMArena Math13071388
MATH Level 556.9%94.7%
FrontierMath (Feb 2025 set)2.1%9.3%
FrontierMath (Tiers 1-3)—14.7%
Omni-MATH27.6%—
FrontierMath Tier 4 (v1)0%—

Knowledge o1 leads

Claude 3.5 Sonnet: 28.6 (#245), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkClaude 3.5 Sonneto1
GPQA Diamond55.3%76.8%
Humanity's Last Exam4.1%8%
Confabulations19.9%11.7%
LMArena Expert12651361
SimpleQA Verified—41.1%
MMLU-Pro77.7%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal o1 leads

Claude 3.5 Sonnet: 26.5 (#120), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkClaude 3.5 Sonneto1
LMArena Vision11251168
GeoBench62%80%
VPCT33%37%
Video-MME60%—
SpatialViz-Bench—41.4%

Multilingual o1 leads

Claude 3.5 Sonnet: 43.2 (#185), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkClaude 3.5 Sonneto1
LMArena Non-English12831358
LMArena Chinese12721394
LMArena French13051344
LMArena German12971337
LMArena Japanese12341346
LMArena Korean12001396
LMArena Russian13061356
LMArena Spanish12901345

Instruction Following o1 leads

Claude 3.5 Sonnet: 68.8 (#182), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkClaude 3.5 Sonneto1
LiveBench Instruction Following69.3%81.5%
LMArena Instruction Following12971367
IFEval85.5%—

Long Context o1 leads

Claude 3.5 Sonnet: 39.9 (#167), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkClaude 3.5 Sonneto1
LMArena Longer Query13111378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Claude 3.5 Sonnet: 52.9 (#164), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkClaude 3.5 Sonneto1
LMArena Text12981366
LMArena Creative Writing12921348
Short-Story Creative Writing80.3%70.2%
LMArena Multi-Turn13261369
LiveBench Language53.8%65.4%
EQ-Bench Creative Writing1451—
WildBench79.2%—

Frequently asked questions

Is Claude 3.5 Sonnet better than o1?

o1 is the stronger model overall, scoring 40.9 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and o1 share?

45 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper