Model comparison

Chatgpt 4o Latest 20250326 vs Claude Haiku 4.5

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 39.5 on the Noometry Index.

Last verified . 17 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 6 categories and Claude Haiku 4.5 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Chatgpt 4o Latest 20250326 leads 33.8 to 15.1.

Side by side

Chatgpt 4o Latest 20250326 and Claude Haiku 4.5 specifications
Chatgpt 4o Latest 20250326Claude Haiku 4.5
ProviderOpenAIAnthropic
Noometry Index43.839.5
Released—2025-10-15
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$1
Output $ / M tokens—$5
Results tracked2153

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Haiku 4.5 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Claude Haiku 4.5: 44.0 (#78)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Coding14131453
SWE-bench Verified (bash only)—66.6%
LMArena WebDev—1330
SWE-bench Multilingual—64.7%
SciCode—43.3%
WeirdML—45.4%
ALE-Bench—653.48

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, Claude Haiku 4.5: 33.6 (#52)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
Terminal-Bench—35.5%
Berkeley Function Calling Leaderboard—68.7%
DeepResearch Bench—45.5%
BALROG—31.2%
ExploitBench—13.7%
Vending-Bench 2—458.89

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Claude Haiku 4.5: 15.1 (#320)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Hard Prompts14241420
ARC-AGI-2—4%
Kagi LLM Benchmark75%—
NYT Connections (extended)—14.3%
ARC-AGI-1—47.7%
CritPt—0%
Chess Puzzles—8%
DTBench—73.6%
LMCA—30.9%
Epoch Capabilities Index—142.41
ForecastBench—61.4

Math Claude Haiku 4.5 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Claude Haiku 4.5: 44.9 (#78)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Math14071396
OTIS Mock AIME 2024-2025—66.7%
Omni-MATH—56.1%
MATH Level 5—96.4%
FrontierMath (Feb 2025 set)—5.9%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Claude Haiku 4.5: 37.7 (#153)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Expert14011442
GPQA Diamond—71.2%
SimpleQA Verified—13.2%
MMLU-Pro—77.7%
Confabulations16.6%—
Vectara Hallucination Rate—9.8%
GPQA (HELM)—60.5%

Multimodal Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Claude Haiku 4.5: 26.8 (#118)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Vision1243—
Blueprint-Bench 2—0%
LMArena Document—1420

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Claude Haiku 4.5: 49.9 (#129)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Non-English14191377
LMArena Chinese14571417
LMArena French14461408
LMArena German14231375
LMArena Japanese14051339
LMArena Korean13961347
LMArena Russian14291381
LMArena Spanish14341420

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Claude Haiku 4.5: 71.4 (#149)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Instruction Following14031414
IFEval—80.1%

Long Context Too close to call

Chatgpt 4o Latest 20250326: 43.1 (#107), Claude Haiku 4.5: 43.6 (#92)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Longer Query14131427

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Claude Haiku 4.5: 57.9 (#123)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Haiku 4.5
LMArena Text14291396
LMArena Creative Writing14051372
LMArena Multi-Turn14541409
EQ-Bench Creative Writing1501—
WildBench—83.9%
EQ-Bench 4—1064

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Claude Haiku 4.5?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 39.5 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Claude Haiku 4.5 better for coding?

Claude Haiku 4.5 scores higher on coding benchmarks: 44.0 versus 41.6 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Claude Haiku 4.5 share?

17 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Claude Haiku 4.5 has 53.

Related comparisons

Go deeper