Model comparison

Claude 3 Opus vs GPT-5 Nano

GPT-5 Nano is the stronger model overall, scoring 33.5 to 29.5 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

GPT-5 Nano OpenAI

33.5

Rank #241 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude 3 Opus scores higher in 2 categories and GPT-5 Nano in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5 Nano leads 29.4 to 14.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.7% for Claude 3 Opus and 81.1% for GPT-5 Nano.

Side by side

Claude 3 Opus and GPT-5 Nano specifications
Claude 3 OpusGPT-5 Nano
ProviderAnthropicOpenAI
Noometry Index29.533.5
Released2024-02-292025-08-07
WeightsProprietaryProprietary
Context window—400K
Max output—128K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.40
Results tracked4649

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3 Opus: 32.9 (#267), GPT-5 Nano: 33.6 (#254)

Coding benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
WeirdML19.2%38.1%
LMArena Coding12641351
SWE-bench Verified (bash only)—34.8%
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
BigCodeBench Complete57.4%—
ALE-Bench—718.67
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use GPT-5 Nano leads

Claude 3 Opus: 24.6 (#116), GPT-5 Nano: 25.8 (#106)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
Terminal-Bench—21.8%
Berkeley Function Calling Leaderboard—51.5%
Cybench10%—
METR Time Horizons29.5%—

Reasoning GPT-5 Nano leads

Claude 3 Opus: 14.6 (#324), GPT-5 Nano: 16.3 (#306)

Reasoning benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
Chess Puzzles5%27%
LMArena Hard Prompts12451328
DTBench61.6%62.7%
LMCA17%7.9%
Epoch Capabilities Index126.91139.38
ForecastBench58.459.1
ARC-AGI-2—2.6%
SimpleBench23.5%—
Kagi LLM Benchmark—62.2%
ARC-AGI-1—20.7%
EnigmaEval0.8%—
LiveBench Reasoning40.6%—
Mystery Game Puzzles—9%
LiveBench Data Analysis57.9%—
LiveBench49.2%—
WinoGrande88.5%—

Math GPT-5 Nano leads

Claude 3 Opus: 14.8 (#299), GPT-5 Nano: 29.4 (#241)

Math benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
OTIS Mock AIME 2024-20254.7%81.1%
LMArena Math12731317
MATH Level 537.5%95.2%
FrontierMath (Tiers 1-3)—20%
FrontierMath Tier 4—2.4%
ProofBench—12%
Omni-MATH—54.6%
LiveBench Math43.6%—
FrontierMath (Feb 2025 set)—8.3%
FrontierMath Tier 4 (v1)—2.1%

Knowledge GPT-5 Nano leads

Claude 3 Opus: 24.5 (#267), GPT-5 Nano: 35.9 (#178)

Knowledge benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
GPQA Diamond47.2%69.4%
SimpleQA Verified12.6%11.7%
LMArena Expert12231321
MMLU-Pro—77.8%
Confabulations22.7%—
Vectara Hallucination Rate—10.5%
GPQA (HELM)—67.9%
MMLU84.6%—

Multimodal GPT-5 Nano leads

Claude 3 Opus: 27.1 (#116), GPT-5 Nano: 31.3 (#108)

Multimodal benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
LMArena Vision10231159
VPCT—37.2%

Multilingual GPT-5 Nano leads

Claude 3 Opus: 41.4 (#207), GPT-5 Nano: 45.3 (#172)

Multilingual benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
LMArena Non-English12581313
LMArena Chinese12481356
LMArena German12581327
LMArena Japanese12041226
LMArena Korean11871269
LMArena Russian12801296
LMArena Spanish12461360
LMArena French1275—

Instruction Following GPT-5 Nano leads

Claude 3 Opus: 64.1 (#228), GPT-5 Nano: 75.0 (#79)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
LMArena Instruction Following12481306
LiveBench Instruction Following63.9%—
IFEval—93.2%

Long Context Claude 3 Opus leads

Claude 3 Opus: 38.2 (#202), GPT-5 Nano: 31.3 (#281)

Long Context benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
LMArena Longer Query12591312
Fiction.LiveBench—44.4%

Writing & Preference Claude 3 Opus leads

Claude 3 Opus: 47.2 (#213), GPT-5 Nano: 39.1 (#249)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGPT-5 Nano
LMArena Text12621320
LMArena Creative Writing12351249
LMArena Multi-Turn12751311
EQ-Bench Creative Writing—705
WildBench—80.6%
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than GPT-5 Nano?

GPT-5 Nano is the stronger model overall, scoring 33.5 to 29.5 on the Noometry Index.

Is Claude 3 Opus or GPT-5 Nano better for coding?

They score almost the same on coding (32.9 vs 33.6); test both on your own repository before choosing.

How many benchmarks do Claude 3 Opus and GPT-5 Nano share?

27 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and GPT-5 Nano has 49.

Related comparisons

Go deeper