Anthropic, proprietary

Claude Sonnet 4

Claude Sonnet 4 by Anthropic ranks 145th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.8. Its strongest category is agentic & tool use, where it ranks 31st. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#145 of 354
Index score
40.8
Evidence
Confirmed 58 results
Provider
Anthropic
Released
May 22, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
64K
Input price
$3 / M
Output price
$15 / M
Blended price
$6 / M
Output speed
31 tokens/s Kagi
Value
#198 of 219
Knowledge cutoff
March 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Sonnet 4 category scores
  1. Coding 43.5
  2. Agentic & Tool Use 38.5
  3. Reasoning 22.9
  4. Math 43.3
  5. Knowledge 41.8
  6. Multimodal 26.2
  7. Multilingual 46.7
  8. Instruction Following 71.7
  9. Long Context 33.7
  10. Writing & Preference 57.1
Claude Sonnet 4 category ranks
CategoryScoreRankResults
Coding43.5#886
Agentic & Tool Use38.5#314
Reasoning22.9#1879
Math43.3#804
Knowledge41.8#1087
Multimodal26.2#1213
Multilingual46.7#1561
Instruction Following71.7#1452
Long Context33.7#2592
Writing & Preference57.1#1326

Strengths and weaknesses

Categories where Claude Sonnet 4 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Sonnet 4: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use38.5+8.1#31 of 154, top 21%
Math43.3+6.8#80 of 327, top 25%
Coding43.5+4.7#88 of 340, top 26%

Weakest categories

Claude Sonnet 4: weakest categories
CategoryScorevs medianRank
Multimodal26.2−12.4#121 of 128, top 95%
Long Context33.7−7.3#259 of 296, top 88%
Reasoning22.9−0.7#187 of 350, top 54%

Closest competitors

The models ranked just above and below Claude Sonnet 4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Sonnet 4
ModelRankScoreBlended $/MSpeed
Grok-3 mini#14141.2—10Compare
Claude Opus 4.1#14241.0$30—Compare
o1#14340.9$26.25—Compare
Gemini 3.1 Flash Lite#14440.8$0.5610Compare
Qwen2.5-Max#14640.7——Compare
Nemotron 3 Nano 30B A3B#14740.6$0.0875—Compare
Granite 4.2 8B#14840.5$0.11—Compare
Step 3#14940.5—7Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Sonnet 4 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)64.9%#17 of 39, top 44%SWE-bench2025-07-26
Aider Polyglot56.4%Epoch AI
Aider Polyglot61.3%#12 of 44, top 28%32KEpoch AI
SciCode40%#79 of 121, top 66%Epoch AI
GSO4.9%#21 of 31, top 68%Epoch AI
WeirdML43.9%Epoch AI
WeirdML46.1%#57 of 119, top 48%16KEpoch AI
LMArena Coding1414#123 of 294, top 42%thinking-32kLMArena2026-10-08
ALE-Bench655.35#74 of 105, top 71%32KEpoch AI

Agentic & Tool Use

Claude Sonnet 4 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
TheAgentCompany33.1%#3 of 14, top 22%Epoch AI
Cybench35%#8 of 21, top 39%Epoch AI
DeepResearch Bench46.6%#13 of 24, top 55%2KEpoch AI
OSWorld43.9%#5 of 8, top 63%Epoch AI
METR Time Horizons62%#17 of 32, top 54%16KEpoch AI

Reasoning

Claude Sonnet 4 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-21.3%Epoch AI
ARC-AGI-25.9%#53 of 83, top 64%16KEpoch AI
ARC-AGI-20.8%1KEpoch AI
ARC-AGI-22.1%8KEpoch AI
SimpleBench45.5%#49 of 77, top 64%Epoch AI
SimpleBench45.5%12KEpoch AI
Kagi LLM Benchmark55.9%Kagi LLM Benchmark
Kagi LLM Benchmark73%#18 of 99, top 19%Kagi LLM Benchmark
ARC-AGI-123.8%Epoch AI
ARC-AGI-140%#61 of 83, top 74%16KEpoch AI
ARC-AGI-128%1KEpoch AI
ARC-AGI-129%8KEpoch AI
CritPt0.3%#92 of 134, top 69%Epoch AI
EnigmaEval3.1%#27 of 38, top 72%Epoch AI
LMArena Hard Prompts1372#144 of 297, top 49%thinking-32kLMArena2026-10-08
DTBench77.1%#81 of 151, top 54%Epoch AI
LMCA29%#79 of 125, top 64%Epoch AI
Epoch Capabilities Index141.69#102 of 213, top 48%Epoch AI2025-05-22
ForecastBench60.2#29 of 72, top 41%Epoch AI

Math

Claude Sonnet 4 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202528.9%Epoch AI2025-05-22
OTIS Mock AIME 2024-202553.3%16KEpoch AI2025-05-22
OTIS Mock AIME 2024-202571.1%#89 of 173, top 52%32KEpoch AI2025-05-22
OTIS Mock AIME 2024-202568.9%59KEpoch AI2025-05-23
Omni-MATH60.2%#10 of 57, top 18%HELM Capabilities
Omni-MATH51.3%HELM Capabilities
LMArena Math1375#146 of 285, top 52%thinking-32kLMArena2026-10-08
MATH Level 584.4%#21 of 79, top 27%Epoch AI2025-05-22
FrontierMath (Feb 2025 set)4.1%#50 of 68, top 74%Epoch AI2025-07-04
FrontierMath Tier 4 (v1)0%#48 of 55, top 88%Epoch AI2025-07-01

Knowledge

Claude Sonnet 4 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond66.7%Epoch AI2025-05-22
GPQA Diamond75.8%16KEpoch AI2025-05-22
GPQA Diamond78.3%32KEpoch AI2025-05-22
GPQA Diamond79.2%#82 of 186, top 45%59KEpoch AI2025-05-26
Humanity's Last Exam7.8%#31 of 41, top 76%Epoch AI
MMLU-Pro84.3%#9 of 58, top 16%HELM Capabilities
MMLU-Pro84.3%#9 of 58, top 16%HELM Capabilities
Confabulations (lower is better)13.2%#10 of 51, top 20%Lech Mazur benchmarks
Confabulations (lower is better)14.8%no reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)10.3%#56 of 96, top 59%Vectara Hallucination Leaderboard
GPQA (HELM)64.3%HELM Capabilities
GPQA (HELM)70.6%#10 of 57, top 18%HELM Capabilities
LMArena Expert1372#140 of 273, top 52%thinking-32kLMArena2026-10-08

Multimodal

Claude Sonnet 4 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1191#84 of 122, top 69%thinking-32kLMArena2026-10-09
GeoBench37%#23 of 25, top 92%Epoch AI
GeoBench32%32KEpoch AI
VPCT30%Epoch AI
VPCT34%#20 of 24, top 84%32KEpoch AI
MindCube44.8%#2 of 2Epoch AI

Multilingual

Claude Sonnet 4 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1333#156 of 297, top 53%thinking-32kLMArena2026-10-08
LMArena Chinese1350#167 of 285, top 59%thinking-32kLMArena2026-10-08
LMArena French1363#140 of 223, top 63%LMArena2026-10-08
LMArena German1331#142 of 231, top 62%LMArena2026-10-08
LMArena Japanese1302#120 of 211, top 57%LMArena2026-10-08
LMArena Korean1291#134 of 213, top 63%LMArena2026-10-08
LMArena Russian1355#144 of 283, top 51%thinking-32kLMArena2026-10-08
LMArena Spanish1357#138 of 226, top 62%thinking-32kLMArena2026-10-08

Instruction Following

Claude Sonnet 4 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval84%#23 of 57, top 41%HELM Capabilities
IFEval83.9%HELM Capabilities
LMArena Instruction Following1376#126 of 298, top 43%thinking-32kLMArena2026-10-08

Long Context

Claude Sonnet 4 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench46.9%#37 of 47, top 79%Epoch AI
LMArena Longer Query1398#119 of 291, top 41%thinking-32kLMArena2026-10-08

Writing & Preference

Claude Sonnet 4 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1351#157 of 297, top 53%thinking-32kLMArena2026-10-08
LMArena Creative Writing1345#134 of 295, top 46%thinking-32kLMArena2026-10-08
Short-Story Creative Writing80.9%Epoch AI
Short-Story Creative Writing81.4%#12 of 39, top 31%16KEpoch AI
EQ-Bench Creative Writing1483#59 of 115, top 52%EQ-Bench
WildBench82.5%HELM Capabilities
WildBench83.8%#16 of 57, top 29%HELM Capabilities
LMArena Multi-Turn1376#137 of 295, top 47%thinking-32kLMArena2026-10-08

API pricing by provider

Claude Sonnet 4 API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$3$15$0.302026-10-10
openrouter$3$15$0.302026-10-10
vertex$3$15$0.302026-10-10

Compare Claude Sonnet 4

Other Anthropic models

Frequently asked questions

How good is Claude Sonnet 4?

Claude Sonnet 4 by Anthropic ranks 145th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.8. Its strongest category is agentic & tool use, where it ranks 31st. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window.

How much does Claude Sonnet 4 cost?

Claude Sonnet 4 costs $3 per million input tokens and $15 per million output tokens on vertex, with cached input at $0.30.

What is Claude Sonnet 4's context window?

Claude Sonnet 4 accepts up to 200K tokens of input and can write up to 64K tokens in one response.

Is Claude Sonnet 4 open source?

No. Claude Sonnet 4 is proprietary and available only through Anthropic's API and partner platforms.

How fast is Claude Sonnet 4?

Claude Sonnet 4 generated about 31 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Claude Sonnet 4's strengths and weaknesses?

Relative to other ranked models, Claude Sonnet 4 places best in agentic & tool use, math, coding and lowest in multimodal, long context, reasoning.

What is Claude Sonnet 4 best at?

Its best category is agentic & tool use, where it ranks 31st on Noometry.