Model comparison

Claude Haiku 4.5 vs GPT-5.2

GPT-5.2 is the stronger model overall, scoring 54.1 to 39.5 on the Noometry Index. Claude Haiku 4.5 costs 2.4× less per token, which makes it the better buy when GPT-5.2's lead doesn't matter for your workload.

Last verified . 41 shared benchmarks.

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

GPT-5.2 OpenAI

54.1

Rank #34 Confirmed

Summary

  • They share 41 benchmarks with published results for both. Claude Haiku 4.5 scores higher in 0 categories and GPT-5.2 in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.2 leads 50.2 to 15.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 14.3% for Claude Haiku 4.5 and 83.6% for GPT-5.2.
  • Claude Haiku 4.5 is cheaper at $1 / $5 per million input/output tokens, against $1.75 / $14 for GPT-5.2.
  • GPT-5.2 accepts more context: 400K tokens versus 200K.

Side by side

Claude Haiku 4.5 and GPT-5.2 specifications
Claude Haiku 4.5GPT-5.2
ProviderAnthropicOpenAI
Noometry Index39.554.1
Released2025-10-152025-12-11
WeightsProprietaryProprietary
Context window200K400K
Max output64K128K
Input $ / M tokens$1$1.75
Output $ / M tokens$5$14
Results tracked5367

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.2 leads

Claude Haiku 4.5: 44.0 (#78), GPT-5.2: 51.6 (#37)

Coding benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
SWE-bench Verified (bash only)66.6%72.8%
LMArena WebDev13301416
SWE-bench Multilingual64.7%66.7%
WeirdML45.4%72.2%
LMArena Coding14531447
ALE-Bench653.481,294
SWE-bench Verified—73.8%
SciCode43.3%—
GSO—27.4%
AlgoTune—2.05

Agentic & Tool Use GPT-5.2 leads

Claude Haiku 4.5: 33.6 (#52), GPT-5.2: 40.2 (#24)

Agentic & Tool Use benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
Terminal-Bench35.5%64.9%
Berkeley Function Calling Leaderboard68.7%55.9%
DeepResearch Bench45.5%41.1%
Vending-Bench 2458.893,591
GDPval—49.7%
Remote Labor Index—2.5%
τ²-bench Airline—83%
τ²-bench Banking—32.2%
τ²-bench Retail—81.6%
τ²-bench Telecom—89.7%
BALROG31.2%—
ExploitBench13.7%—
LMArena Search—1207
METR Time Horizons—75.3%

Reasoning GPT-5.2 leads

Claude Haiku 4.5: 15.1 (#320), GPT-5.2: 50.2 (#35)

Reasoning benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
ARC-AGI-24%52.9%
NYT Connections (extended)14.3%83.6%
ARC-AGI-147.7%86.2%
Chess Puzzles8%49%
LMArena Hard Prompts14201445
DTBench73.6%90.9%
LMCA30.9%43.9%
Epoch Capabilities Index142.41153.45
ForecastBench61.460.1
SimpleBench—45.8%
Kagi LLM Benchmark—73.3%
CritPt0%—
EnigmaEval—10.4%
EBR-Bench—23%
Mystery Game Puzzles—23%

Math GPT-5.2 leads

Claude Haiku 4.5: 44.9 (#78), GPT-5.2: 60.0 (#38)

Knowledge GPT-5.2 leads

Claude Haiku 4.5: 37.7 (#153), GPT-5.2: 59.3 (#32)

Knowledge benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
GPQA Diamond71.2%91.4%
SimpleQA Verified13.2%37.1%
Vectara Hallucination Rate9.8%8.4%
LMArena Expert14421445
Humanity's Last Exam—27.8%
MMLU-Pro77.7%—
GPQA (HELM)60.5%—

Multimodal GPT-5.2 leads

Claude Haiku 4.5: 26.8 (#118), GPT-5.2: 51.3 (#7)

Multimodal benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
LMArena Document14201405
LMArena Vision—1268
VPCT—84%
Blueprint-Bench 20%—
Furniture Assembly—38.3%

Multilingual GPT-5.2 leads

Claude Haiku 4.5: 49.9 (#129), GPT-5.2: 53.4 (#67)

Multilingual benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
LMArena Non-English13771425
LMArena Chinese14171460
LMArena French14081455
LMArena German13751448
LMArena Japanese13391420
LMArena Korean13471392
LMArena Russian13811440
LMArena Spanish14201433

Instruction Following GPT-5.2 leads

Claude Haiku 4.5: 71.4 (#149), GPT-5.2: 74.7 (#89)

Instruction Following benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
LMArena Instruction Following14141417
IFEval80.1%—

Long Context Too close to call

Claude Haiku 4.5: 43.6 (#92), GPT-5.2: 44.0 (#78)

Long Context benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
LMArena Longer Query14271428
CL-bench—18.2%

Writing & Preference GPT-5.2 leads

Claude Haiku 4.5: 57.9 (#123), GPT-5.2: 66.8 (#32)

Writing & Preference benchmarks
BenchmarkClaude Haiku 4.5GPT-5.2
LMArena Text13961439
LMArena Creative Writing13721401
LMArena Multi-Turn14091458
EQ-Bench Creative Writing—1703
WildBench83.9%—
EQ-Bench 41064—

Frequently asked questions

Is Claude Haiku 4.5 better than GPT-5.2?

GPT-5.2 is the stronger model overall, scoring 54.1 to 39.5 on the Noometry Index. Claude Haiku 4.5 costs 2.4× less per token, which makes it the better buy when GPT-5.2's lead doesn't matter for your workload.

Which is cheaper, Claude Haiku 4.5 or GPT-5.2?

Claude Haiku 4.5 is cheaper. It lists at $1 per million input tokens and $5 per million output tokens; GPT-5.2 lists at $1.75 and $14.

Is Claude Haiku 4.5 or GPT-5.2 better for coding?

GPT-5.2 scores higher on coding benchmarks: 51.6 versus 44.0 in the Noometry coding category.

Which has the bigger context window?

GPT-5.2 does, with 400K tokens against 200K.

How many benchmarks do Claude Haiku 4.5 and GPT-5.2 share?

41 benchmarks have published results for both models. Claude Haiku 4.5 has 53 scored results on Noometry and GPT-5.2 has 67.

Related comparisons

Go deeper