Model comparison

Claude Opus 4 vs GPT-5 Mini

Claude Opus 4 is the stronger model overall, scoring 43.1 to 41.8 on the Noometry Index. GPT-5 Mini costs 44× less per token, which makes it the better buy when Claude Opus 4's lead doesn't matter for your workload.

Last verified . 48 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Summary

  • They share 48 benchmarks with published results for both. Claude Opus 4 scores higher in 5 categories and GPT-5 Mini in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Claude Opus 4 leads 47.2 to 40.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 64.4% for Claude Opus 4 and 86.7% for GPT-5 Mini.
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • GPT-5 Mini accepts more context: 400K tokens versus 200K.

Side by side

Claude Opus 4 and GPT-5 Mini specifications
Claude Opus 4GPT-5 Mini
ProviderAnthropicOpenAI
Noometry Index43.141.8
Released2025-05-222025-08-07
WeightsProprietaryProprietary
Context window200K400K
Max output32K128K
Input $ / M tokens$15$0.25
Output $ / M tokens$75$2
Results tracked5660

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), GPT-5 Mini: 40.1 (#146)

Coding benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
SWE-bench Verified70.7%64.7%
SWE-bench Verified (bash only)67.6%59.8%
WeirdML43.7%52.7%
LMArena Coding14421406
AlgoTune1.331.38
Aider Polyglot72%—
SWE-bench Multilingual—39.7%
SciCode—39.2%
GSO6.9%—
ALE-Bench—799.77

Agentic & Tool Use Claude Opus 4 leads

Claude Opus 4: 34.8 (#42), GPT-5 Mini: 31.1 (#70)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
Terminal-Bench—34.8%
Berkeley Function Calling Leaderboard—55.5%
Cybench38%—
DeepResearch Bench46.8%—
LMArena Search1127—
METR Time Horizons63.9%—
Vending-Bench 2—-31.18

Reasoning Claude Opus 4 leads

Claude Opus 4: 27.3 (#121), GPT-5 Mini: 23.9 (#168)

Reasoning benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
ARC-AGI-28.6%4.4%
Kagi LLM Benchmark74.3%70.3%
ARC-AGI-135.7%54.3%
CritPt0.3%0%
EnigmaEval5.6%8.2%
LMArena Hard Prompts13991380
DTBench81.6%80.5%
LMCA37.4%34.2%
Epoch Capabilities Index142.67145.52
ForecastBench61.161
SimpleBench58.8%—
Chess Puzzles—30%
Mystery Game Puzzles—10%

Math GPT-5 Mini leads

Claude Opus 4: 42.0 (#86), GPT-5 Mini: 46.7 (#69)

Math benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
OTIS Mock AIME 2024-202564.4%86.7%
Omni-MATH61.6%72.2%
LMArena Math13901378
MATH Level 585%97.8%
FrontierMath (Feb 2025 set)4.5%27.2%
FrontierMath Tier 4 (v1)4.2%6.3%
FrontierMath (Tiers 1-3)—46.7%
FrontierMath Tier 4—12.2%
ProofBench—9%

Knowledge GPT-5 Mini leads

Claude Opus 4: 44.0 (#88), GPT-5 Mini: 45.6 (#86)

Knowledge benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
GPQA Diamond76.3%75%
Humanity's Last Exam10.7%19.4%
MMLU-Pro87.5%83.5%
Confabulations15.9%13.3%
Vectara Hallucination Rate12%12.9%
GPQA (HELM)70.8%75.6%
LMArena Expert13861379
SimpleQA Verified—21.6%

Multimodal GPT-5 Mini leads

Claude Opus 4: 31.5 (#106), GPT-5 Mini: 35.6 (#85)

Multimodal benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
LMArena Vision11921202
VPCT38%40.2%
GeoBench49%—

Multilingual Too close to call

Claude Opus 4: 48.8 (#138), GPT-5 Mini: 48.9 (#137)

Multilingual benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
LMArena Non-English13621363
LMArena Chinese13861385
LMArena French13721386
LMArena German13911366
LMArena Japanese13311341
LMArena Korean13211308
LMArena Russian13921362
LMArena Spanish13891355

Instruction Following Too close to call

Claude Opus 4: 77.1 (#28), GPT-5 Mini: 76.2 (#46)

Instruction Following benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
IFEval91.8%92.7%
LMArena Instruction Following14061357

Long Context GPT-5 Mini leads

Claude Opus 4: 39.6 (#172), GPT-5 Mini: 41.9 (#132)

Long Context benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
Fiction.LiveBench61.1%69.4%
LMArena Longer Query14221355

Writing & Preference Claude Opus 4 leads

Claude Opus 4: 61.2 (#89), GPT-5 Mini: 55.2 (#148)

Writing & Preference benchmarks
BenchmarkClaude Opus 4GPT-5 Mini
LMArena Text13771373
LMArena Creative Writing13871325
Short-Story Creative Writing83.6%83.1%
EQ-Bench Creative Writing15801313
WildBench85.2%85.5%
LMArena Multi-Turn13961363

Frequently asked questions

Is Claude Opus 4 better than GPT-5 Mini?

Claude Opus 4 is the stronger model overall, scoring 43.1 to 41.8 on the Noometry Index. GPT-5 Mini costs 44× less per token, which makes it the better buy when Claude Opus 4's lead doesn't matter for your workload.

Which is cheaper, Claude Opus 4 or GPT-5 Mini?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or GPT-5 Mini better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

GPT-5 Mini does, with 400K tokens against 200K.

How many benchmarks do Claude Opus 4 and GPT-5 Mini share?

48 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and GPT-5 Mini has 60.

Related comparisons

Go deeper