Model comparison

Claude Opus 4 vs Claude Sonnet 4

Claude Opus 4 is the stronger model overall, scoring 43.1 to 40.8 on the Noometry Index. Claude Sonnet 4 costs 5.0× less per token, which makes it the better buy when Claude Opus 4's lead doesn't matter for your workload.

Last verified . 53 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Summary

  • They share 53 benchmarks with published results for both. Claude Opus 4 scores higher in 8 categories and Claude Sonnet 4 in 2 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Claude Opus 4 leads 39.6 to 33.7.
  • The biggest single-benchmark swing is Fiction.LiveBench: 61.1% for Claude Opus 4 and 46.9% for Claude Sonnet 4.
  • Claude Sonnet 4 is cheaper at $3 / $15 per million input/output tokens, against $15 / $75 for Claude Opus 4.

Side by side

Claude Opus 4 and Claude Sonnet 4 specifications
Claude Opus 4Claude Sonnet 4
ProviderAnthropicAnthropic
Noometry Index43.140.8
Released2025-05-222025-05-22
WeightsProprietaryProprietary
Context window200K200K
Max output32K64K
Input $ / M tokens$15$3
Output $ / M tokens$75$15
Results tracked5658

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), Claude Sonnet 4: 43.5 (#88)

Coding benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
SWE-bench Verified (bash only)67.6%64.9%
Aider Polyglot72%61.3%
GSO6.9%4.9%
WeirdML43.7%46.1%
LMArena Coding14421414
SWE-bench Verified70.7%—
SciCode—40%
ALE-Bench—655.35
AlgoTune1.33—

Agentic & Tool Use Claude Sonnet 4 leads

Claude Opus 4: 34.8 (#42), Claude Sonnet 4: 38.5 (#31)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
Cybench38%35%
DeepResearch Bench46.8%46.6%
METR Time Horizons63.9%62%
TheAgentCompany—33.1%
OSWorld—43.9%
LMArena Search1127—

Reasoning Claude Opus 4 leads

Claude Opus 4: 27.3 (#121), Claude Sonnet 4: 22.9 (#187)

Reasoning benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
ARC-AGI-28.6%5.9%
SimpleBench58.8%45.5%
Kagi LLM Benchmark74.3%73%
ARC-AGI-135.7%40%
CritPt0.3%0.3%
EnigmaEval5.6%3.1%
LMArena Hard Prompts13991372
DTBench81.6%77.1%
LMCA37.4%29%
Epoch Capabilities Index142.67141.69
ForecastBench61.160.2

Math Claude Sonnet 4 leads

Claude Opus 4: 42.0 (#86), Claude Sonnet 4: 43.3 (#80)

Math benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
OTIS Mock AIME 2024-202564.4%71.1%
Omni-MATH61.6%60.2%
LMArena Math13901375
MATH Level 585%84.4%
FrontierMath (Feb 2025 set)4.5%4.1%
FrontierMath Tier 4 (v1)4.2%0%

Knowledge Claude Opus 4 leads

Claude Opus 4: 44.0 (#88), Claude Sonnet 4: 41.8 (#108)

Knowledge benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
GPQA Diamond76.3%79.2%
Humanity's Last Exam10.7%7.8%
MMLU-Pro87.5%84.3%
Confabulations15.9%13.2%
Vectara Hallucination Rate12%10.3%
GPQA (HELM)70.8%70.6%
LMArena Expert13861372

Multimodal Claude Opus 4 leads

Claude Opus 4: 31.5 (#106), Claude Sonnet 4: 26.2 (#121)

Multimodal benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
LMArena Vision11921191
GeoBench49%37%
VPCT38%34%
MindCube—44.8%

Multilingual Claude Opus 4 leads

Claude Opus 4: 48.8 (#138), Claude Sonnet 4: 46.7 (#156)

Multilingual benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
LMArena Non-English13621333
LMArena Chinese13861350
LMArena French13721363
LMArena German13911331
LMArena Japanese13311302
LMArena Korean13211291
LMArena Russian13921355
LMArena Spanish13891357

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), Claude Sonnet 4: 71.7 (#145)

Instruction Following benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
IFEval91.8%84%
LMArena Instruction Following14061376

Long Context Claude Opus 4 leads

Claude Opus 4: 39.6 (#172), Claude Sonnet 4: 33.7 (#259)

Long Context benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
Fiction.LiveBench61.1%46.9%
LMArena Longer Query14221398

Writing & Preference Claude Opus 4 leads

Claude Opus 4: 61.2 (#89), Claude Sonnet 4: 57.1 (#132)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Claude Sonnet 4
LMArena Text13771351
LMArena Creative Writing13871345
Short-Story Creative Writing83.6%81.4%
EQ-Bench Creative Writing15801483
WildBench85.2%83.8%
LMArena Multi-Turn13961376

Frequently asked questions

Is Claude Opus 4 better than Claude Sonnet 4?

Claude Opus 4 is the stronger model overall, scoring 43.1 to 40.8 on the Noometry Index. Claude Sonnet 4 costs 5.0× less per token, which makes it the better buy when Claude Opus 4's lead doesn't matter for your workload.

Which is cheaper, Claude Opus 4 or Claude Sonnet 4?

Claude Sonnet 4 is cheaper. It lists at $3 per million input tokens and $15 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Claude Sonnet 4 better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 43.5 in the Noometry coding category.

Which has the bigger context window?

Both accept 200K tokens.

How many benchmarks do Claude Opus 4 and Claude Sonnet 4 share?

53 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Claude Sonnet 4 has 58.

Related comparisons

Go deeper