Model comparison

Claude 3.5 Haiku vs o1

o1 is the stronger model overall, scoring 40.9 to 29.2 on the Noometry Index.

Last verified . 37 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 1 category and o1 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o1 leads 41.5 to 18.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 73.3% for o1.

Side by side

Claude 3.5 Haiku and o1 specifications
Claude 3.5 Haikuo1
ProviderAnthropicOpenAI
Noometry Index29.240.9
Released2024-10-222024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked4952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Claude 3.5 Haiku: 32.9 (#265), o1: 46.1 (#70)

Coding benchmarks
BenchmarkClaude 3.5 Haikuo1
Aider Polyglot28%61.7%
WeirdML30.7%47.6%
LiveBench Coding51.4%69.7%
LMArena Coding12861367
CadEval32%56%
SciCode27.4%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete59%—
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Claude 3.5 Haiku leads

Claude 3.5 Haiku: 28.0 (#95), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 Haikuo1
Cybench—10%
BALROG19.3%—
METR Time Horizons—51.1%

Reasoning o1 leads

Claude 3.5 Haiku: 17.7 (#290), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkClaude 3.5 Haikuo1
LiveBench Reasoning28.1%91.6%
LMArena Hard Prompts12511371
DTBench56.7%74.7%
LiveBench Data Analysis48.5%65.5%
Epoch Capabilities Index127.15141.91
LiveBench43.5%75.7%
SimpleBench—41.7%
ARC-AGI-1—30.7%
CritPt0%—
Chess Puzzles—15%
EnigmaEval—5.7%
LMCA—22.3%

Math o1 leads

Claude 3.5 Haiku: 14.7 (#300), o1: 36.1 (#175)

Math benchmarks
BenchmarkClaude 3.5 Haikuo1
OTIS Mock AIME 2024-20254.3%73.3%
LiveBench Math35.5%80.3%
LMArena Math12441388
MATH Level 546.4%94.7%
FrontierMath (Feb 2025 set)0.3%9.3%
FrontierMath (Tiers 1-3)—14.7%
Omni-MATH22.4%—

Knowledge o1 leads

Claude 3.5 Haiku: 18.7 (#281), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkClaude 3.5 Haikuo1
GPQA Diamond38.1%76.8%
Confabulations36.7%11.7%
LMArena Expert12081361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
MMLU-Pro60.5%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal o1 leads

Claude 3.5 Haiku: 26.8 (#117), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkClaude 3.5 Haikuo1
LMArena Vision10921168
GeoBench34%80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Claude 3.5 Haiku: 40.0 (#218), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkClaude 3.5 Haikuo1
LMArena Non-English12381358
LMArena Chinese12291394
LMArena French12641344
LMArena German12371337
LMArena Japanese11751346
LMArena Korean11731396
LMArena Russian12531356
LMArena Spanish12611345

Instruction Following o1 leads

Claude 3.5 Haiku: 62.9 (#234), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkClaude 3.5 Haikuo1
LiveBench Instruction Following61.9%81.5%
LMArena Instruction Following12411367
IFEval79.2%—

Long Context o1 leads

Claude 3.5 Haiku: 38.3 (#200), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkClaude 3.5 Haikuo1
LMArena Longer Query12611378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Claude 3.5 Haiku: 42.7 (#234), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkClaude 3.5 Haikuo1
LMArena Text12551366
LMArena Creative Writing12331348
Short-Story Creative Writing73.5%70.2%
LMArena Multi-Turn12651369
LiveBench Language35.4%65.4%
EQ-Bench Creative Writing1146—
WildBench76%—

Frequently asked questions

Is Claude 3.5 Haiku better than o1?

o1 is the stronger model overall, scoring 40.9 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and o1 share?

37 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper