Anthropic, proprietary

Claude Sonnet 4.5

Claude Sonnet 4.5 by Anthropic ranks 81st of 354 ranked models on the Noometry Index as of October 2026, with a score of 44.1. Its strongest category is agentic & tool use, where it ranks 32nd. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#81 of 354
Index score
44.1
Evidence
Confirmed 73 results
Provider
Anthropic
Released
September 29, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
64K
Input price
$3 / M
Output price
$15 / M
Blended price
$6 / M
Output speed
85 tokens/s Kagi
Value
#196 of 219
Knowledge cutoff
July 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Claude Sonnet 4.5 category scores
  1. Coding 47.3
  2. Agentic & Tool Use 38.3
  3. Reasoning 26.9
  4. Math 32.3
  5. Knowledge 48.4
  6. Multimodal 34.8
  7. Multilingual 53.4
  8. Instruction Following 75.0
  9. Long Context 45.2
  10. Writing & Preference 66.5
Claude Sonnet 4.5 category ranks
CategoryScoreRankResults
Coding47.3#618
Agentic & Tool Use38.3#3211
Reasoning26.9#12513
Math32.3#2167
Knowledge48.4#767
Multimodal34.8#891
Multilingual53.4#691
Instruction Following75.0#782
Long Context45.2#461
Writing & Preference66.5#345

Strengths and weaknesses

Categories where Claude Sonnet 4.5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude Sonnet 4.5: strongest categories
CategoryScorevs medianRank
Writing & Preference66.5+12.8#34 of 312, top 11%
Long Context45.2+4.3#46 of 296, top 16%
Coding47.3+8.6#61 of 340, top 18%

Weakest categories

Claude Sonnet 4.5: weakest categories
CategoryScorevs medianRank
Multimodal34.8−3.8#89 of 128, top 70%
Math32.3−4.3#216 of 327, top 67%
Reasoning26.9+3.3#125 of 350, top 36%

Closest competitors

The models ranked just above and below Claude Sonnet 4.5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude Sonnet 4.5
ModelRankScoreBlended $/MSpeed
Amazon Nova Experimental Chat 26 02 10#7744.5——Compare
DeepSeek-V3.2-Exp#7844.3$0.2916Compare
Hy3#7944.2$0.14—Compare
Inkling#8044.1$2.57—Compare
Chatgpt 4o Latest 20250326#8243.8—21Compare
ERNIE 5.1#8343.8——Compare
GLM-5V-Turbo#8443.8$1.90—Compare
MiniMax-M3#8543.8$0.52—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude Sonnet 4.5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified71.3%#23 of 32, top 72%Epoch AI2026-02-05
SWE-bench Verified (bash only)71.4%#9 of 39, top 24%highSWE-bench2026-02-17
LMArena WebDev1385LMArena2026-10-08
LMArena WebDev1393#76 of 113, top 68%LMArena2026-10-08
SWE-bench Multilingual67%#8 of 13, top 62%SWE-bench2026-02-13
SciCode44.7%#61 of 121, top 51%Epoch AI
GSO14.7%#15 of 31, top 49%Epoch AI
WeirdML46.7%Epoch AI
WeirdML47.7%#53 of 119, top 45%16KEpoch AI
LMArena Coding1485LMArena2026-10-08
LMArena Coding1489#31 of 294, top 11%LMArena2026-10-08
ALE-Bench796.15#59 of 105, top 57%32KEpoch AI
AlgoTune1.52#9 of 18, top 50%Epoch AI

Agentic & Tool Use

Claude Sonnet 4.5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench46.5%#19 of 41, top 47%Epoch AI
Berkeley Function Calling Leaderboard73.2%#2 of 49, top 5%fcBerkeley Function Calling Leaderboard
GDPval42.5%#4 of 11, top 37%Epoch AI
Remote Labor Index2.1%#11 of 14, top 79%Epoch AI
τ²-bench Airline72%#7 of 7, top 100%enabledτ²-bench2026-02-26
τ²-bench Banking25.3%#17 of 26, top 66%enabledτ²-bench2026-02-26
τ²-bench Retail72.4%#7 of 7, top 100%enabledτ²-bench2026-02-26
τ²-bench Telecom84.9%#7 of 7, top 100%enabledτ²-bench2026-02-26
Cybench60%#3 of 21, top 15%Epoch AI
DeepResearch Bench52.6%#5 of 24, top 21%2KEpoch AI
DeepResearch Bench47.5%lowEpoch AI
OSWorld62.9%#4 of 8, top 50%Epoch AI
LMArena Search1159#23 of 32, top 72%LMArena2026-08-24
METR Time Horizons67.4%#11 of 32, top 35%16KEpoch AI
Vending-Bench 23,839#37 of 60, top 62%Epoch AI

Reasoning

Claude Sonnet 4.5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-23.8%Epoch AI
ARC-AGI-26.9%16KEpoch AI
ARC-AGI-25.8%1KEpoch AI
ARC-AGI-213.6%#46 of 83, top 56%32KEpoch AI
ARC-AGI-26.9%8KEpoch AI
SimpleBench54.3%#35 of 77, top 46%Epoch AI
SimpleBench54.3%12KEpoch AI
Kagi LLM Benchmark57.9%#44 of 99, top 45%Kagi LLM Benchmark
NYT Connections (extended)37.3%#69 of 91, top 76%Lech Mazur benchmarks
NYT Connections (extended)35.8%no reasoningLech Mazur benchmarks
ARC-AGI-125.5%Epoch AI
ARC-AGI-148.3%16KEpoch AI
ARC-AGI-131%1KEpoch AI
ARC-AGI-163.7%#47 of 83, top 57%32KEpoch AI
ARC-AGI-146.5%8KEpoch AI
CritPt1.1%#78 of 134, top 59%Epoch AI
Chess Puzzles4%Epoch AI2026-08-06
Chess Puzzles12%#76 of 129, top 59%32KEpoch AI2025-12-08
EnigmaEval6%#20 of 38, top 53%Epoch AI
EBR-Bench2.4%#23 of 24, top 96%Epoch AI2026-06-25
LMArena Hard Prompts1462#45 of 297, top 16%LMArena2026-10-08
LMArena Hard Prompts1462LMArena2026-10-08
Mystery Game Puzzles17%#50 of 74, top 68%48KEpoch AI2026-07-25
DTBench83.2%#59 of 151, top 40%Epoch AI
LMCA38.8%#48 of 125, top 39%Epoch AI
Epoch Capabilities Index146.84#68 of 213, top 32%Epoch AI2025-09-29
ForecastBench61.9#4 of 72, top 6%Epoch AI

Math

Claude Sonnet 4.5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)23.9%#66 of 81, top 82%32KEpoch AI2026-06-11
FrontierMath Tier 42.4%#58 of 63, top 93%32KEpoch AI2026-06-11
OTIS Mock AIME 2024-202535.6%Epoch AI2025-09-29
OTIS Mock AIME 2024-202571.1%16KEpoch AI2025-10-28
OTIS Mock AIME 2024-202577.8%#82 of 173, top 48%32KEpoch AI2025-10-21
OTIS Mock AIME 2024-202577.8%59KEpoch AI2025-10-28
ProofBench19%#50 of 77, top 65%Epoch AI
Omni-MATH55.3%#14 of 57, top 25%HELM Capabilities
LMArena Math1422LMArena2026-10-08
LMArena Math1449#63 of 285, top 23%LMArena2026-10-08
MATH Level 597.7%#5 of 79, top 7%32KEpoch AI2025-10-21
FrontierMath (Feb 2025 set)9.3%Epoch AI2025-11-16
FrontierMath (Feb 2025 set)15.2%#34 of 68, top 50%32KEpoch AI2025-11-16
FrontierMath (Feb 2025 set)13.5%59KEpoch AI2025-11-13
FrontierMath Tier 4 (v1)2.1%Epoch AI2025-09-29
FrontierMath Tier 4 (v1)4.2%#30 of 55, top 55%32KEpoch AI2025-10-22

Knowledge

Claude Sonnet 4.5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond73.7%Epoch AI2025-09-29
GPQA Diamond78.8%16KEpoch AI2025-10-28
GPQA Diamond81.7%32KEpoch AI2025-10-21
GPQA Diamond82.3%#74 of 186, top 40%59KEpoch AI2025-10-28
Humanity's Last Exam13.7%#21 of 41, top 52%Epoch AI
SimpleQA Verified23.7%Epoch AI2026-08-10
SimpleQA Verified30.7%#57 of 77, top 75%59KEpoch AI2026-08-27
MMLU-Pro86.9%#3 of 58, top 6%HELM Capabilities
Vectara Hallucination Rate (lower is better)12%#74 of 96, top 78%Vectara Hallucination Leaderboard
GPQA (HELM)68.6%#11 of 57, top 20%HELM Capabilities
LMArena Expert1472LMArena2026-10-08
LMArena Expert1482#40 of 273, top 15%LMArena2026-10-08

Multimodal

Claude Sonnet 4.5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
VPCT38%Epoch AI
VPCT39.8%#14 of 24, top 59%32KEpoch AI
LMArena Document1450#23 of 38, top 61%LMArena2026-09-13

Multilingual

Claude Sonnet 4.5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1418LMArena2026-10-08
LMArena Non-English1425#69 of 297, top 24%LMArena2026-10-08
LMArena Chinese1459#93 of 285, top 33%LMArena2026-10-08
LMArena Chinese1459LMArena2026-10-08
LMArena French1458#54 of 223, top 25%LMArena2026-10-08
LMArena French1451LMArena2026-10-08
LMArena German1427#69 of 231, top 30%LMArena2026-10-08
LMArena German1421LMArena2026-10-08
LMArena Japanese1390#69 of 211, top 33%LMArena2026-10-08
LMArena Japanese1373LMArena2026-10-08
LMArena Korean1367LMArena2026-10-08
LMArena Korean1403#51 of 213, top 24%LMArena2026-10-08
LMArena Russian1429LMArena2026-10-08
LMArena Russian1437#54 of 283, top 20%LMArena2026-10-08
LMArena Spanish1457#39 of 226, top 18%LMArena2026-10-08
LMArena Spanish1444LMArena2026-10-08

Instruction Following

Claude Sonnet 4.5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval85%#19 of 57, top 34%HELM Capabilities
LMArena Instruction Following1459#35 of 298, top 12%LMArena2026-10-08
LMArena Instruction Following1458LMArena2026-10-08

Long Context

Claude Sonnet 4.5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1475LMArena2026-10-08
LMArena Longer Query1476#28 of 291, top 10%LMArena2026-10-08

Writing & Preference

Claude Sonnet 4.5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1439#66 of 297, top 23%LMArena2026-10-08
LMArena Text1435LMArena2026-10-08
LMArena Creative Writing1425LMArena2026-10-08
LMArena Creative Writing1442#35 of 295, top 12%LMArena2026-10-08
EQ-Bench Creative Writing1678#34 of 115, top 30%EQ-Bench
WildBench85.4%#9 of 57, top 16%HELM Capabilities
LMArena Multi-Turn1465#35 of 295, top 12%LMArena2026-10-08
LMArena Multi-Turn1454LMArena2026-10-08

API pricing by provider

Claude Sonnet 4.5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
anthropic$3$15$0.302026-10-10
azure$3$15$0.302026-10-10
bedrock$3$15$0.302026-10-10
openrouter$3$15$0.302026-10-10
vertex$3$15$0.302026-10-10

Compare Claude Sonnet 4.5

Other Anthropic models

Frequently asked questions

How good is Claude Sonnet 4.5?

Claude Sonnet 4.5 by Anthropic ranks 81st of 354 ranked models on the Noometry Index as of October 2026, with a score of 44.1. Its strongest category is agentic & tool use, where it ranks 32nd. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window.

How much does Claude Sonnet 4.5 cost?

Claude Sonnet 4.5 costs $3 per million input tokens and $15 per million output tokens on Anthropic's own API, with cached input at $0.30.

What is Claude Sonnet 4.5's context window?

Claude Sonnet 4.5 accepts up to 200K tokens of input and can write up to 64K tokens in one response.

Is Claude Sonnet 4.5 open source?

No. Claude Sonnet 4.5 is proprietary and available only through Anthropic's API and partner platforms.

How fast is Claude Sonnet 4.5?

Claude Sonnet 4.5 generated about 85 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Claude Sonnet 4.5's strengths and weaknesses?

Relative to other ranked models, Claude Sonnet 4.5 places best in writing & preference, long context, coding and lowest in multimodal, math, reasoning.

What is Claude Sonnet 4.5 best at?

Its best category is agentic & tool use, where it ranks 32nd on Noometry.