Model comparison

Claude Opus 4 vs GPT-4

Claude Opus 4 is the stronger model overall, scoring 43.1 to 29.1 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude Opus 4 scores higher in 8 categories and GPT-4 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 4 leads 42.0 to 10.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 64.4% for Claude Opus 4 and 1.1% for GPT-4.
  • Claude Opus 4 is cheaper at $15 / $75 per million input/output tokens, against $30 / $60 for GPT-4.
  • Claude Opus 4 accepts more context: 200K tokens versus 8K.

Side by side

Claude Opus 4 and GPT-4 specifications
Claude Opus 4GPT-4
ProviderAnthropicOpenAI
Noometry Index43.129.1
Released2025-05-222023-03-14
WeightsProprietaryProprietary
Context window200K8K
Max output32K8K
Input $ / M tokens$15$30
Output $ / M tokens$75$60
Results tracked5638

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkClaude Opus 4GPT-4
WeirdML43.7%12.4%
LMArena Coding14421254
SWE-bench Verified70.7%—
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
GSO6.9%—
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
AlgoTune1.33—
HumanEval+—79.3%

Agentic & Tool Use Not comparable

Claude Opus 4: 34.8 (#42), GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4GPT-4
METR Time Horizons63.9%36.1%
Cybench38%—
DeepResearch Bench46.8%—
LMArena Search1127—

Reasoning Claude Opus 4 leads

Claude Opus 4: 27.3 (#121), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkClaude Opus 4GPT-4
LMArena Hard Prompts13991241
DTBench81.6%62.7%
LMCA37.4%17.1%
Epoch Capabilities Index142.67125.89
ForecastBench61.157.8
ARC-AGI-28.6%—
SimpleBench58.8%—
Kagi LLM Benchmark74.3%—
ARC-AGI-135.7%—
CritPt0.3%—
Chess Puzzles—4%
EnigmaEval5.6%—
Mystery Game Puzzles—12%
BIG-Bench Hard—75.1%
HellaSwag—95.3%
WinoGrande—87.5%

Math Claude Opus 4 leads

Claude Opus 4: 42.0 (#86), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkClaude Opus 4GPT-4
OTIS Mock AIME 2024-202564.4%1.1%
LMArena Math13901269
MATH Level 585%23%
Omni-MATH61.6%—
FrontierMath (Feb 2025 set)4.5%—
FrontierMath Tier 4 (v1)4.2%—
GSM8K—92%

Knowledge Claude Opus 4 leads

Claude Opus 4: 44.0 (#88), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkClaude Opus 4GPT-4
GPQA Diamond76.3%35.7%
LMArena Expert13861211
Humanity's Last Exam10.7%—
MMLU-Pro87.5%—
Confabulations15.9%—
Vectara Hallucination Rate12%—
GPQA (HELM)70.8%—
MMLU—86.4%
TriviaQA—84.8%

Multimodal Not comparable

Claude Opus 4: 31.5 (#106), GPT-4: —

Multimodal benchmarks
BenchmarkClaude Opus 4GPT-4
LMArena Vision1192—
GeoBench49%—
VPCT38%—

Multilingual Claude Opus 4 leads

Claude Opus 4: 48.8 (#138), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkClaude Opus 4GPT-4
LMArena Non-English13621246
LMArena Chinese13861242
LMArena French13721283
LMArena German13911251
LMArena Japanese13311209
LMArena Korean13211184
LMArena Russian13921251
LMArena Spanish13891261

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkClaude Opus 4GPT-4
LMArena Instruction Following14061241
IFEval91.8%—

Long Context Claude Opus 4 leads

Claude Opus 4: 39.6 (#172), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkClaude Opus 4GPT-4
LMArena Longer Query14221244
Fiction.LiveBench61.1%—

Writing & Preference Claude Opus 4 leads

Claude Opus 4: 61.2 (#89), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkClaude Opus 4GPT-4
LMArena Text13771263
LMArena Creative Writing13871244
EQ-Bench Creative Writing1580752
LMArena Multi-Turn13961257
Short-Story Creative Writing83.6%—
WildBench85.2%—

Frequently asked questions

Is Claude Opus 4 better than GPT-4?

Claude Opus 4 is the stronger model overall, scoring 43.1 to 29.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or GPT-4?

Claude Opus 4 is cheaper. It lists at $15 per million input tokens and $75 per million output tokens; GPT-4 lists at $30 and $60.

Is Claude Opus 4 or GPT-4 better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 31.6 in the Noometry coding category.

Which has the bigger context window?

Claude Opus 4 does, with 200K tokens against 8K.

How many benchmarks do Claude Opus 4 and GPT-4 share?

27 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper