Model comparison

Claude 2.1 vs GPT-5 Pro

GPT-5 Pro is the stronger model overall, scoring 46.4 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

GPT-5 Pro OpenAI

46.4

Rank #64 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and GPT-5 Pro in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5 Pro leads 56.7 to 15.4.
  • The biggest single-benchmark swing is WeirdML: 7.1% for Claude 2.1 and 60.4% for GPT-5 Pro.

Side by side

Claude 2.1 and GPT-5 Pro specifications
Claude 2.1GPT-5 Pro
ProviderAnthropicOpenAI
Noometry Index25.246.4
Released2023-11-212025-10-06
WeightsProprietaryProprietary
Context window—400K
Max output—272K
Input $ / M tokens—$15
Output $ / M tokens—$120
Results tracked712

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 Pro leads

Claude 2.1: 26.2 (#327), GPT-5 Pro: 44.0 (#80)

Coding benchmarks
BenchmarkClaude 2.1GPT-5 Pro
WeirdML7.1%60.4%
AlgoTune—1.31

Reasoning GPT-5 Pro leads

Claude 2.1: 21.4 (#221), GPT-5 Pro: 38.9 (#62)

Reasoning benchmarks
BenchmarkClaude 2.1GPT-5 Pro
Epoch Capabilities Index119.27150.28
ARC-AGI-2—18.3%
SimpleBench—61.6%
Kagi LLM Benchmark—76.8%
ARC-AGI-1—70.2%
EnigmaEval—18.8%
DTBench51%—
ForecastBench54.2—

Math GPT-5 Pro leads

Claude 2.1: 10.2 (#315), GPT-5 Pro: 48.5 (#63)

Math benchmarks
BenchmarkClaude 2.1GPT-5 Pro
FrontierMath (Tiers 1-3)—55.8%
FrontierMath Tier 4—19.5%
OTIS Mock AIME 2024-20251.9%—
FrontierMath Tier 4 (v1)—14.6%

Knowledge GPT-5 Pro leads

Claude 2.1: 15.4 (#292), GPT-5 Pro: 56.7 (#42)

Knowledge benchmarks
BenchmarkClaude 2.1GPT-5 Pro
GPQA Diamond33%—
Humanity's Last Exam—31.6%
MMLU73.5%—

Frequently asked questions

Is Claude 2.1 better than GPT-5 Pro?

GPT-5 Pro is the stronger model overall, scoring 46.4 to 25.2 on the Noometry Index.

Is Claude 2.1 or GPT-5 Pro better for coding?

GPT-5 Pro scores higher on coding benchmarks: 44.0 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and GPT-5 Pro share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and GPT-5 Pro has 12.

Related comparisons

Go deeper