Model comparison

Claude 3.5 Sonnet vs o1-mini

Claude 3.5 Sonnet and o1-mini score almost the same on the Noometry Index (34.6 vs 34.0), so choose on price, context window or the category you care about most.

Last verified . 37 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 5 categories and o1-mini in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1-mini leads 35.4 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 46.9% for o1-mini.

Side by side

Claude 3.5 Sonnet and o1-mini specifications
Claude 3.5 Sonneto1-mini
ProviderAnthropicOpenAI
Noometry Index34.634.0
Released2024-06-202024-09-12
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6039

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
Aider Polyglot51.6%32.9%
WeirdML40%36.3%
LiveBench Coding67.1%48%
LMArena Coding13421362
HumanEval+81.7%89%
MBPP+74.3%78.8%
GSO4.6%—
BigCodeBench Instruct46.8%—
BigCodeBench Complete58.6%—
CadEval48%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
Cybench17.5%10%
TheAgentCompany24%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
SimpleBench41.4%18.1%
LiveBench Reasoning56.7%72.3%
LMArena Hard Prompts13051333
LiveBench Data Analysis55%57.9%
Epoch Capabilities Index133.55135.82
LiveBench59%57.8%
ARC-AGI-2—0.8%
ARC-AGI-1—14%
EnigmaEval0.9%—
DTBench67.8%—
ForecastBench60.7—

Math o1-mini leads

Claude 3.5 Sonnet: 19.2 (#288), o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
OTIS Mock AIME 2024-20258.5%46.9%
LiveBench Math52.3%62%
LMArena Math13071358
MATH Level 556.9%89.2%
FrontierMath (Feb 2025 set)2.1%1.7%
Omni-MATH27.6%—
FrontierMath Tier 4 (v1)0%—

Knowledge o1-mini leads

Claude 3.5 Sonnet: 28.6 (#245), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
GPQA Diamond55.3%62.4%
Confabulations19.9%18.6%
LMArena Expert12651316
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), o1-mini: —

Multimodal benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Too close to call

Claude 3.5 Sonnet: 43.2 (#185), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
LMArena Non-English12831289
LMArena Chinese12721314
LMArena French13051293
LMArena German12971278
LMArena Japanese12341245
LMArena Korean12001223
LMArena Russian13061283
LMArena Spanish12901303

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
LiveBench Instruction Following69.3%65.4%
LMArena Instruction Following12971304
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
LMArena Longer Query13111320

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkClaude 3.5 Sonneto1-mini
LMArena Text12981317
LMArena Creative Writing12921244
Short-Story Creative Writing80.3%64.9%
LMArena Multi-Turn13261314
LiveBench Language53.8%40.9%
EQ-Bench Creative Writing1451—
WildBench79.2%—

Frequently asked questions

Is Claude 3.5 Sonnet better than o1-mini?

Claude 3.5 Sonnet and o1-mini score almost the same on the Noometry Index (34.6 vs 34.0), so choose on price, context window or the category you care about most.

Is Claude 3.5 Sonnet or o1-mini better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 35.5 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and o1-mini share?

37 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper