Model comparison

Claude 3 Opus vs GPT-4.1

GPT-4.1 is the stronger model overall, scoring 35.9 to 29.5 on the Noometry Index.

Last verified . 30 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Claude 3 Opus scores higher in 1 category and GPT-4.1 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-4.1 leads 37.1 to 24.5.
  • The biggest single-benchmark swing is MATH Level 5: 37.5% for Claude 3 Opus and 83% for GPT-4.1.

Side by side

Claude 3 Opus and GPT-4.1 specifications
Claude 3 OpusGPT-4.1
ProviderAnthropicOpenAI
Noometry Index29.535.9
Released2024-02-292025-04-14
WeightsProprietaryProprietary
Context window—1.05M
Max output—33K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked4652

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.1 leads

Claude 3 Opus: 32.9 (#267), GPT-4.1: 34.4 (#238)

Coding benchmarks
BenchmarkClaude 3 OpusGPT-4.1
WeirdML19.2%39%
LMArena Coding12641391
SWE-bench Verified—48.5%
SWE-bench Verified (bash only)—39.6%
Aider Polyglot—52.4%
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
BigCodeBench Complete57.4%—
CadEval—42%
ALE-Bench—558.1
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use GPT-4.1 leads

Claude 3 Opus: 24.6 (#116), GPT-4.1: 34.7 (#43)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGPT-4.1
Berkeley Function Calling Leaderboard—54%
Cybench10%—
METR Time Horizons29.5%—

Reasoning Claude 3 Opus leads

Claude 3 Opus: 14.6 (#324), GPT-4.1: 11.7 (#339)

Reasoning benchmarks
BenchmarkClaude 3 OpusGPT-4.1
SimpleBench23.5%27%
Chess Puzzles5%6%
EnigmaEval0.8%2.2%
LMArena Hard Prompts12451384
DTBench61.6%68.3%
LMCA17%25.6%
Epoch Capabilities Index126.91136.78
ForecastBench58.461.5
ARC-AGI-2—0.4%
Kagi LLM Benchmark—52.3%
ARC-AGI-1—5.5%
LiveBench Reasoning40.6%—
LiveBench Data Analysis57.9%—
LiveBench49.2%—
WinoGrande88.5%—

Math GPT-4.1 leads

Claude 3 Opus: 14.8 (#299), GPT-4.1: 22.3 (#280)

Math benchmarks
BenchmarkClaude 3 OpusGPT-4.1
OTIS Mock AIME 2024-20254.7%38.3%
LMArena Math12731370
MATH Level 537.5%83%
FrontierMath (Tiers 1-3)—6%
Omni-MATH—47.1%
LiveBench Math43.6%—
FrontierMath (Feb 2025 set)—5.5%
FrontierMath Tier 4 (v1)—0%

Knowledge GPT-4.1 leads

Claude 3 Opus: 24.5 (#267), GPT-4.1: 37.1 (#160)

Knowledge benchmarks
BenchmarkClaude 3 OpusGPT-4.1
GPQA Diamond47.2%66.9%
SimpleQA Verified12.6%31.1%
LMArena Expert12231364
Humanity's Last Exam—5.4%
MMLU-Pro—81.1%
Confabulations22.7%—
Vectara Hallucination Rate—5.6%
GPQA (HELM)—65.9%
MMLU84.6%—

Multimodal GPT-4.1 leads

Claude 3 Opus: 27.1 (#116), GPT-4.1: 38.2 (#67)

Multimodal benchmarks
BenchmarkClaude 3 OpusGPT-4.1
LMArena Vision10231211
GeoBench—72%

Multilingual GPT-4.1 leads

Claude 3 Opus: 41.4 (#207), GPT-4.1: 49.4 (#133)

Multilingual benchmarks
BenchmarkClaude 3 OpusGPT-4.1
LMArena Non-English12581370
LMArena Chinese12481382
LMArena French12751382
LMArena German12581381
LMArena Japanese12041319
LMArena Korean11871339
LMArena Russian12801377
LMArena Spanish12461376

Instruction Following GPT-4.1 leads

Claude 3 Opus: 64.1 (#228), GPT-4.1: 71.3 (#153)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGPT-4.1
LMArena Instruction Following12481367
LiveBench Instruction Following63.9%—
IFEval—83.8%

Long Context GPT-4.1 leads

Claude 3 Opus: 38.2 (#202), GPT-4.1: 40.0 (#163)

Long Context benchmarks
BenchmarkClaude 3 OpusGPT-4.1
LMArena Longer Query12591385
Fiction.LiveBench—63.9%

Writing & Preference GPT-4.1 leads

Claude 3 Opus: 47.2 (#213), GPT-4.1: 57.6 (#125)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGPT-4.1
LMArena Text12621383
LMArena Creative Writing12351363
LMArena Multi-Turn12751398
EQ-Bench Creative Writing—1420
WildBench—85.4%
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than GPT-4.1?

GPT-4.1 is the stronger model overall, scoring 35.9 to 29.5 on the Noometry Index.

Is Claude 3 Opus or GPT-4.1 better for coding?

GPT-4.1 scores higher on coding benchmarks: 34.4 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and GPT-4.1 share?

30 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and GPT-4.1 has 52.

Related comparisons

Go deeper