# Chatgpt 4o Latest 20250326 vs Claude 3.5 Haiku

> Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 29.2 on the Noometry Index.

- Canonical page: https://noometry.com/compare/chatgpt-4o-vs-claude-3-5-haiku
- Last updated: 2026-10-10
- Shared benchmarks: 20

## Summary

- They share 20 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 9 categories and Claude 3.5 Haiku in 0 categories; 9 gaps are clear of the uncertainty.
- The widest gap is in math, where Chatgpt 4o Latest 20250326 leads 38.6 to 14.7.
- The biggest single-benchmark swing is Confabulations: 16.6% for Chatgpt 4o Latest 20250326 and 36.7% for Claude 3.5 Haiku.

## Snapshot

| | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| Provider | OpenAI | Anthropic |
| Noometry Index | 43.8 | 29.2 |
| Rank | 82 | 315 |
| Context | — | — |
| Input $/M | — | — |
| Output $/M | — | — |
| Weights | Proprietary | Proprietary |

## Coding

- Chatgpt 4o Latest 20250326: 41.6 (#122)
- Claude 3.5 Haiku: 32.9 (#265)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Coding | 1413 | 1286 |
| Aider Polyglot | — | 28% |
| SciCode | — | 27.4% |
| WeirdML | — | 30.7% |
| BigCodeBench Instruct | — | 46.1% |
| LiveBench Coding | — | 51.4% |
| BigCodeBench Complete | — | 59% |
| CadEval | — | 32% |

## Agentic & Tool Use

- Chatgpt 4o Latest 20250326: —
- Claude 3.5 Haiku: 28.0 (#95)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| BALROG | — | 19.3% |

## Reasoning

- Chatgpt 4o Latest 20250326: 33.8 (#71)
- Claude 3.5 Haiku: 17.7 (#290)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Hard Prompts | 1424 | 1251 |
| Kagi LLM Benchmark | 75% | — |
| CritPt | — | 0% |
| LiveBench Reasoning | — | 28.1% |
| DTBench | — | 56.7% |
| LiveBench Data Analysis | — | 48.5% |
| Epoch Capabilities Index | — | 127.15 |
| LiveBench | — | 43.5% |

## Math

- Chatgpt 4o Latest 20250326: 38.6 (#134)
- Claude 3.5 Haiku: 14.7 (#300)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Math | 1407 | 1244 |
| OTIS Mock AIME 2024-2025 | — | 4.3% |
| Omni-MATH | — | 22.4% |
| LiveBench Math | — | 35.5% |
| MATH Level 5 | — | 46.4% |
| FrontierMath (Feb 2025 set) | — | 0.3% |

## Knowledge

- Chatgpt 4o Latest 20250326: 39.2 (#137)
- Claude 3.5 Haiku: 18.7 (#281)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| Confabulations | 16.6% | 36.7% |
| LMArena Expert | 1401 | 1208 |
| GPQA Diamond | — | 38.1% |
| MMLU-Pro | — | 60.5% |
| GPQA (HELM) | — | 36.3% |
| MMLU | — | 74.3% |

## Multimodal

- Chatgpt 4o Latest 20250326: 39.6 (#58)
- Claude 3.5 Haiku: 26.8 (#117)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Vision | 1243 | 1092 |
| GeoBench | — | 34% |

## Multilingual

- Chatgpt 4o Latest 20250326: 52.9 (#76)
- Claude 3.5 Haiku: 40.0 (#218)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Non-English | 1419 | 1238 |
| LMArena Chinese | 1457 | 1229 |
| LMArena French | 1446 | 1264 |
| LMArena German | 1423 | 1237 |
| LMArena Japanese | 1405 | 1175 |
| LMArena Korean | 1396 | 1173 |
| LMArena Russian | 1429 | 1253 |
| LMArena Spanish | 1434 | 1261 |

## Instruction Following

- Chatgpt 4o Latest 20250326: 74.0 (#107)
- Claude 3.5 Haiku: 62.9 (#234)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Instruction Following | 1403 | 1241 |
| LiveBench Instruction Following | — | 61.9% |
| IFEval | — | 79.2% |

## Long Context

- Chatgpt 4o Latest 20250326: 43.1 (#107)
- Claude 3.5 Haiku: 38.3 (#200)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Longer Query | 1413 | 1261 |

## Writing & Preference

- Chatgpt 4o Latest 20250326: 62.6 (#72)
- Claude 3.5 Haiku: 42.7 (#234)

| Benchmark | Chatgpt 4o Latest 20250326 | Claude 3.5 Haiku |
|---|---|---|
| LMArena Text | 1429 | 1255 |
| LMArena Creative Writing | 1405 | 1233 |
| EQ-Bench Creative Writing | 1501 | 1146 |
| LMArena Multi-Turn | 1454 | 1265 |
| Short-Story Creative Writing | — | 73.5% |
| WildBench | — | 76% |
| LiveBench Language | — | 35.4% |

## FAQ

### Is Chatgpt 4o Latest 20250326 better than Claude 3.5 Haiku?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 29.2 on the Noometry Index.

### Is Chatgpt 4o Latest 20250326 or Claude 3.5 Haiku better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 32.9 in the Noometry coding category.

### How many benchmarks do Chatgpt 4o Latest 20250326 and Claude 3.5 Haiku share?

20 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Claude 3.5 Haiku has 49.
