Model comparison

Claude Opus 4 vs GPT-5.2

GPT-5.2 is the stronger model overall, scoring 54.1 to 43.1 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

GPT-5.2 OpenAI

54.1

Rank #34 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Opus 4 scores higher in 1 category and GPT-5.2 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.2 leads 50.2 to 27.3.
  • The biggest single-benchmark swing is ARC-AGI-1: 35.7% for Claude Opus 4 and 86.2% for GPT-5.2.
  • GPT-5.2 is cheaper at $1.75 / $14 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • GPT-5.2 accepts more context: 400K tokens versus 200K.

Side by side

Claude Opus 4 and GPT-5.2 specifications
Claude Opus 4GPT-5.2
ProviderAnthropicOpenAI
Noometry Index43.154.1
Released2025-05-222025-12-11
WeightsProprietaryProprietary
Context window200K400K
Max output32K128K
Input $ / M tokens$15$1.75
Output $ / M tokens$75$14
Results tracked5667

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.2 leads

Claude Opus 4: 47.2 (#62), GPT-5.2: 51.6 (#37)

Coding benchmarks
BenchmarkClaude Opus 4GPT-5.2
SWE-bench Verified70.7%73.8%
SWE-bench Verified (bash only)67.6%72.8%
GSO6.9%27.4%
WeirdML43.7%72.2%
LMArena Coding14421447
AlgoTune1.332.05
Aider Polyglot72%—
LMArena WebDev—1416
SWE-bench Multilingual—66.7%
ALE-Bench—1,294

Agentic & Tool Use GPT-5.2 leads

Claude Opus 4: 34.8 (#42), GPT-5.2: 40.2 (#24)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4GPT-5.2
DeepResearch Bench46.8%41.1%
LMArena Search11271207
METR Time Horizons63.9%75.3%
Terminal-Bench—64.9%
Berkeley Function Calling Leaderboard—55.9%
GDPval—49.7%
Remote Labor Index—2.5%
τ²-bench Airline—83%
τ²-bench Banking—32.2%
τ²-bench Retail—81.6%
τ²-bench Telecom—89.7%
Cybench38%—
Vending-Bench 2—3,591

Reasoning GPT-5.2 leads

Claude Opus 4: 27.3 (#121), GPT-5.2: 50.2 (#35)

Reasoning benchmarks
BenchmarkClaude Opus 4GPT-5.2
ARC-AGI-28.6%52.9%
SimpleBench58.8%45.8%
Kagi LLM Benchmark74.3%73.3%
ARC-AGI-135.7%86.2%
EnigmaEval5.6%10.4%
LMArena Hard Prompts13991445
DTBench81.6%90.9%
LMCA37.4%43.9%
Epoch Capabilities Index142.67153.45
ForecastBench61.160.1
NYT Connections (extended)—83.6%
CritPt0.3%—
Chess Puzzles—49%
EBR-Bench—23%
Mystery Game Puzzles—23%

Math GPT-5.2 leads

Claude Opus 4: 42.0 (#86), GPT-5.2: 60.0 (#38)

Knowledge GPT-5.2 leads

Claude Opus 4: 44.0 (#88), GPT-5.2: 59.3 (#32)

Knowledge benchmarks
BenchmarkClaude Opus 4GPT-5.2
GPQA Diamond76.3%91.4%
Humanity's Last Exam10.7%27.8%
Vectara Hallucination Rate12%8.4%
LMArena Expert13861445
SimpleQA Verified—37.1%
MMLU-Pro87.5%—
Confabulations15.9%—
GPQA (HELM)70.8%—

Multimodal GPT-5.2 leads

Claude Opus 4: 31.5 (#106), GPT-5.2: 51.3 (#7)

Multimodal benchmarks
BenchmarkClaude Opus 4GPT-5.2
LMArena Vision11921268
VPCT38%84%
GeoBench49%—
Furniture Assembly—38.3%
LMArena Document—1405

Multilingual GPT-5.2 leads

Claude Opus 4: 48.8 (#138), GPT-5.2: 53.4 (#67)

Multilingual benchmarks
BenchmarkClaude Opus 4GPT-5.2
LMArena Non-English13621425
LMArena Chinese13861460
LMArena French13721455
LMArena German13911448
LMArena Japanese13311420
LMArena Korean13211392
LMArena Russian13921440
LMArena Spanish13891433

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), GPT-5.2: 74.7 (#89)

Instruction Following benchmarks
BenchmarkClaude Opus 4GPT-5.2
LMArena Instruction Following14061417
IFEval91.8%—

Long Context GPT-5.2 leads

Claude Opus 4: 39.6 (#172), GPT-5.2: 44.0 (#78)

Long Context benchmarks
BenchmarkClaude Opus 4GPT-5.2
LMArena Longer Query14221428
Fiction.LiveBench61.1%—
CL-bench—18.2%

Writing & Preference GPT-5.2 leads

Claude Opus 4: 61.2 (#89), GPT-5.2: 66.8 (#32)

Writing & Preference benchmarks
BenchmarkClaude Opus 4GPT-5.2
LMArena Text13771439
LMArena Creative Writing13871401
EQ-Bench Creative Writing15801703
LMArena Multi-Turn13961458
Short-Story Creative Writing83.6%—
WildBench85.2%—

Frequently asked questions

Is Claude Opus 4 better than GPT-5.2?

GPT-5.2 is the stronger model overall, scoring 54.1 to 43.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or GPT-5.2?

GPT-5.2 is cheaper. It lists at $1.75 per million input tokens and $14 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or GPT-5.2 better for coding?

GPT-5.2 scores higher on coding benchmarks: 51.6 versus 47.2 in the Noometry coding category.

Which has the bigger context window?

GPT-5.2 does, with 400K tokens against 200K.

How many benchmarks do Claude Opus 4 and GPT-5.2 share?

43 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and GPT-5.2 has 67.

Related comparisons

Go deeper