Model comparison

Claude 3.5 Haiku vs o3

o3 is the stronger model overall, scoring 47.5 to 29.2 on the Noometry Index.

Last verified . 37 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and o3 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 18.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 84.4% for o3.

Side by side

Claude 3.5 Haiku and o3 specifications
Claude 3.5 Haikuo3
ProviderAnthropicOpenAI
Noometry Index29.247.5
Released2024-10-222025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked4963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Claude 3.5 Haiku: 32.9 (#265), o3: 46.8 (#64)

Coding benchmarks
BenchmarkClaude 3.5 Haikuo3
Aider Polyglot28%81.3%
WeirdML30.7%52.4%
LMArena Coding12861408
CadEval32%74%
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
SciCode27.4%—
GSO—8.8%
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
ALE-Bench—933.55

Agentic & Tool Use o3 leads

Claude 3.5 Haiku: 28.0 (#95), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 Haikuo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
BALROG19.3%—
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Claude 3.5 Haiku: 17.7 (#290), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkClaude 3.5 Haikuo3
CritPt0%1.4%
LMArena Hard Prompts12511402
DTBench56.7%84.8%
Epoch Capabilities Index127.15146.86
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
Chess Puzzles—38%
EnigmaEval—13.1%
LiveBench Reasoning28.1%—
Mystery Game Puzzles—29%
LiveBench Data Analysis48.5%—
LMCA—39.7%
ForecastBench—62.5
LiveBench43.5%—

Math o3 leads

Claude 3.5 Haiku: 14.7 (#300), o3: 50.2 (#58)

Math benchmarks
BenchmarkClaude 3.5 Haikuo3
OTIS Mock AIME 2024-20254.3%84.4%
Omni-MATH22.4%71.4%
LMArena Math12441426
MATH Level 546.4%97.8%
FrontierMath (Feb 2025 set)0.3%18.7%
FrontierMath (Tiers 1-3)—33.3%
LiveBench Math35.5%—
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Claude 3.5 Haiku: 18.7 (#281), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkClaude 3.5 Haikuo3
GPQA Diamond38.1%81.8%
MMLU-Pro60.5%85.9%
Confabulations36.7%14.4%
GPQA (HELM)36.3%75.3%
LMArena Expert12081402
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU74.3%—

Multimodal o3 leads

Claude 3.5 Haiku: 26.8 (#117), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkClaude 3.5 Haikuo3
LMArena Vision10921214
GeoBench34%74%
VPCT—52%

Multilingual o3 leads

Claude 3.5 Haiku: 40.0 (#218), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkClaude 3.5 Haikuo3
LMArena Non-English12381401
LMArena Chinese12291437
LMArena French12641430
LMArena German12371420
LMArena Japanese11751403
LMArena Korean11731370
LMArena Russian12531406
LMArena Spanish12611395

Instruction Following o3 leads

Claude 3.5 Haiku: 62.9 (#234), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkClaude 3.5 Haikuo3
IFEval79.2%86.9%
LMArena Instruction Following12411368
LiveBench Instruction Following61.9%—

Long Context o3 leads

Claude 3.5 Haiku: 38.3 (#200), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkClaude 3.5 Haikuo3
LMArena Longer Query12611372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Claude 3.5 Haiku: 42.7 (#234), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkClaude 3.5 Haikuo3
LMArena Text12551410
LMArena Creative Writing12331359
Short-Story Creative Writing73.5%83.9%
EQ-Bench Creative Writing11461676
WildBench76%86.1%
LMArena Multi-Turn12651405
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than o3?

o3 is the stronger model overall, scoring 47.5 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and o3 share?

37 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper