Model comparison

Claude Opus 4.5 vs GPT-5.2

GPT-5.2 is the stronger model overall, scoring 54.1 to 50.5 on the Noometry Index.

Last verified . 65 shared benchmarks.

Claude Opus 4.5 Anthropic

50.5

Rank #47 Confirmed

GPT-5.2 OpenAI

54.1

Rank #34 Confirmed

Summary

  • They share 65 benchmarks with published results for both. Claude Opus 4.5 scores higher in 6 categories and GPT-5.2 in 4 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.2 leads 60.0 to 38.6.
  • The biggest single-benchmark swing is VPCT: 40% for Claude Opus 4.5 and 84% for GPT-5.2.
  • GPT-5.2 is cheaper at $1.75 / $14 per million input/output tokens, against $5 / $25 for Claude Opus 4.5.
  • GPT-5.2 accepts more context: 400K tokens versus 200K.

Side by side

Claude Opus 4.5 and GPT-5.2 specifications
Claude Opus 4.5GPT-5.2
ProviderAnthropicOpenAI
Noometry Index50.554.1
Released2025-11-012025-12-11
WeightsProprietaryProprietary
Context window200K400K
Max output64K128K
Input $ / M tokens$5$1.75
Output $ / M tokens$25$14
Results tracked6967

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.5 leads

Claude Opus 4.5: 54.8 (#27), GPT-5.2: 51.6 (#37)

Coding benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
SWE-bench Verified76.7%73.8%
SWE-bench Verified (bash only)76.8%72.8%
LMArena WebDev14941416
SWE-bench Multilingual70.7%66.7%
GSO26.5%27.4%
WeirdML63.7%72.2%
LMArena Coding15041447
ALE-Bench1,0251,294
AlgoTune1.772.05

Agentic & Tool Use Claude Opus 4.5 leads

Claude Opus 4.5: 47.3 (#12), GPT-5.2: 40.2 (#24)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
Terminal-Bench63.1%64.9%
Berkeley Function Calling Leaderboard77.5%55.9%
GDPval45.5%49.7%
Remote Labor Index3.8%2.5%
τ²-bench Airline84%83%
τ²-bench Banking24.7%32.2%
τ²-bench Retail79.6%81.6%
τ²-bench Telecom92.3%89.7%
DeepResearch Bench54.8%41.1%
LMArena Search11801207
METR Time Horizons75%75.3%
Vending-Bench 24,9673,591
Cybench82%—
OSWorld66.3%—
BALROG43.5%—

Reasoning GPT-5.2 leads

Claude Opus 4.5: 42.6 (#51), GPT-5.2: 50.2 (#35)

Reasoning benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
ARC-AGI-237.6%52.9%
SimpleBench62%45.8%
Kagi LLM Benchmark80.2%73.3%
NYT Connections (extended)52.5%83.6%
ARC-AGI-180%86.2%
Chess Puzzles12%49%
EnigmaEval11.9%10.4%
EBR-Bench14.3%23%
LMArena Hard Prompts14761445
Mystery Game Puzzles22%23%
DTBench89.9%90.9%
LMCA44.5%43.9%
Epoch Capabilities Index150.09153.45
ForecastBench60.760.1

Math GPT-5.2 leads

Claude Opus 4.5: 38.6 (#132), GPT-5.2: 60.0 (#38)

Knowledge GPT-5.2 leads

Claude Opus 4.5: 56.5 (#44), GPT-5.2: 59.3 (#32)

Knowledge benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
GPQA Diamond86%91.4%
Humanity's Last Exam25.2%27.8%
SimpleQA Verified45.7%37.1%
Vectara Hallucination Rate10.9%8.4%
LMArena Expert14871445

Multimodal GPT-5.2 leads

Claude Opus 4.5: 31.4 (#107), GPT-5.2: 51.3 (#7)

Multimodal benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
VPCT40%84%
Furniture Assembly28.3%38.3%
LMArena Document14621405
LMArena Vision—1268
GeoBench75%—

Multilingual Too close to call

Claude Opus 4.5: 54.3 (#47), GPT-5.2: 53.4 (#67)

Multilingual benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
LMArena Non-English14381425
LMArena Chinese14701460
LMArena French14711455
LMArena German14491448
LMArena Japanese14161420
LMArena Korean14241392
LMArena Russian14471440
LMArena Spanish14581433

Instruction Following Claude Opus 4.5 leads

Claude Opus 4.5: 77.5 (#19), GPT-5.2: 74.7 (#89)

Instruction Following benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
LMArena Instruction Following14781417

Long Context Claude Opus 4.5 leads

Claude Opus 4.5: 46.5 (#22), GPT-5.2: 44.0 (#78)

Long Context benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
CL-bench21.1%18.2%
LMArena Longer Query14801428

Writing & Preference Claude Opus 4.5 leads

Claude Opus 4.5: 68.1 (#28), GPT-5.2: 66.8 (#32)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.5GPT-5.2
LMArena Text14511439
LMArena Creative Writing14451401
EQ-Bench Creative Writing16871703
LMArena Multi-Turn14661458

Frequently asked questions

Is Claude Opus 4.5 better than GPT-5.2?

GPT-5.2 is the stronger model overall, scoring 54.1 to 50.5 on the Noometry Index.

Which is cheaper, Claude Opus 4.5 or GPT-5.2?

GPT-5.2 is cheaper. It lists at $1.75 per million input tokens and $14 per million output tokens; Claude Opus 4.5 lists at $5 and $25.

Is Claude Opus 4.5 or GPT-5.2 better for coding?

Claude Opus 4.5 scores higher on coding benchmarks: 54.8 versus 51.6 in the Noometry coding category.

Which has the bigger context window?

GPT-5.2 does, with 400K tokens against 200K.

How many benchmarks do Claude Opus 4.5 and GPT-5.2 share?

65 benchmarks have published results for both models. Claude Opus 4.5 has 69 scored results on Noometry and GPT-5.2 has 67.

Related comparisons

Go deeper