Model comparison

Claude Sonnet 4 vs GPT-5 Mini

GPT-5 Mini is the stronger model overall, scoring 41.8 to 40.8 on the Noometry Index.

Last verified . 48 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Summary

  • They share 48 benchmarks with published results for both. Claude Sonnet 4 scores higher in 3 categories and GPT-5 Mini in 7 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in multimodal, where GPT-5 Mini leads 35.6 to 26.2.
  • The biggest single-benchmark swing is Fiction.LiveBench: 46.9% for Claude Sonnet 4 and 69.4% for GPT-5 Mini.
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • GPT-5 Mini accepts more context: 400K tokens versus 200K.

Side by side

Claude Sonnet 4 and GPT-5 Mini specifications
Claude Sonnet 4GPT-5 Mini
ProviderAnthropicOpenAI
Noometry Index40.841.8
Released2025-05-222025-08-07
WeightsProprietaryProprietary
Context window200K400K
Max output64K128K
Input $ / M tokens$3$0.25
Output $ / M tokens$15$2
Results tracked5860

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude Sonnet 4: 43.5 (#88), GPT-5 Mini: 40.1 (#146)

Coding benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
SWE-bench Verified (bash only)64.9%59.8%
SciCode40%39.2%
WeirdML46.1%52.7%
LMArena Coding14141406
ALE-Bench655.35799.77
SWE-bench Verified—64.7%
Aider Polyglot61.3%—
SWE-bench Multilingual—39.7%
GSO4.9%—
AlgoTune—1.38

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), GPT-5 Mini: 31.1 (#70)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
Terminal-Bench—34.8%
Berkeley Function Calling Leaderboard—55.5%
TheAgentCompany33.1%—
Cybench35%—
DeepResearch Bench46.6%—
OSWorld43.9%—
METR Time Horizons62%—
Vending-Bench 2—-31.18

Reasoning GPT-5 Mini leads

Claude Sonnet 4: 22.9 (#187), GPT-5 Mini: 23.9 (#168)

Reasoning benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
ARC-AGI-25.9%4.4%
Kagi LLM Benchmark73%70.3%
ARC-AGI-140%54.3%
CritPt0.3%0%
EnigmaEval3.1%8.2%
LMArena Hard Prompts13721380
DTBench77.1%80.5%
LMCA29%34.2%
Epoch Capabilities Index141.69145.52
ForecastBench60.261
SimpleBench45.5%—
Chess Puzzles—30%
Mystery Game Puzzles—10%

Math GPT-5 Mini leads

Claude Sonnet 4: 43.3 (#80), GPT-5 Mini: 46.7 (#69)

Math benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
OTIS Mock AIME 2024-202571.1%86.7%
Omni-MATH60.2%72.2%
LMArena Math13751378
MATH Level 584.4%97.8%
FrontierMath (Feb 2025 set)4.1%27.2%
FrontierMath Tier 4 (v1)0%6.3%
FrontierMath (Tiers 1-3)—46.7%
FrontierMath Tier 4—12.2%
ProofBench—9%

Knowledge GPT-5 Mini leads

Claude Sonnet 4: 41.8 (#108), GPT-5 Mini: 45.6 (#86)

Knowledge benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
GPQA Diamond79.2%75%
Humanity's Last Exam7.8%19.4%
MMLU-Pro84.3%83.5%
Confabulations13.2%13.3%
Vectara Hallucination Rate10.3%12.9%
GPQA (HELM)70.6%75.6%
LMArena Expert13721379
SimpleQA Verified—21.6%

Multimodal GPT-5 Mini leads

Claude Sonnet 4: 26.2 (#121), GPT-5 Mini: 35.6 (#85)

Multimodal benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
LMArena Vision11911202
VPCT34%40.2%
GeoBench37%—
MindCube44.8%—

Multilingual GPT-5 Mini leads

Claude Sonnet 4: 46.7 (#156), GPT-5 Mini: 48.9 (#137)

Multilingual benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
LMArena Non-English13331363
LMArena Chinese13501385
LMArena French13631386
LMArena German13311366
LMArena Japanese13021341
LMArena Korean12911308
LMArena Russian13551362
LMArena Spanish13571355

Instruction Following GPT-5 Mini leads

Claude Sonnet 4: 71.7 (#145), GPT-5 Mini: 76.2 (#46)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
IFEval84%92.7%
LMArena Instruction Following13761357

Long Context GPT-5 Mini leads

Claude Sonnet 4: 33.7 (#259), GPT-5 Mini: 41.9 (#132)

Long Context benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
Fiction.LiveBench46.9%69.4%
LMArena Longer Query13981355

Writing & Preference Claude Sonnet 4 leads

Claude Sonnet 4: 57.1 (#132), GPT-5 Mini: 55.2 (#148)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4GPT-5 Mini
LMArena Text13511373
LMArena Creative Writing13451325
Short-Story Creative Writing81.4%83.1%
EQ-Bench Creative Writing14831313
WildBench83.8%85.5%
LMArena Multi-Turn13761363

Frequently asked questions

Is Claude Sonnet 4 better than GPT-5 Mini?

GPT-5 Mini is the stronger model overall, scoring 41.8 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or GPT-5 Mini?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or GPT-5 Mini better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

GPT-5 Mini does, with 400K tokens against 200K.

How many benchmarks do Claude Sonnet 4 and GPT-5 Mini share?

48 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and GPT-5 Mini has 60.

Related comparisons

Go deeper