Anthropic, proprietary

Claude 3.5 Sonnet

Claude 3.5 Sonnet by Anthropic ranks 231st of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.6. Its strongest category is agentic & tool use, where it ranks 67th.

Last verified

Specifications

Noometry rank
#231 of 354
Index score
34.6
Evidence
Confirmed 60 results
Provider
Anthropic
Released
June 20, 2024
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Claude 3.5 Sonnet category scores
  1. Coding 39.0
  2. Agentic & Tool Use 32.3
  3. Reasoning 23.1
  4. Math 19.2
  5. Knowledge 28.6
  6. Multimodal 26.5
  7. Multilingual 43.2
  8. Instruction Following 68.8
  9. Long Context 39.9
  10. Writing & Preference 52.9
Claude 3.5 Sonnet category ranks
CategoryScoreRankResults
Coding39.0#1658
Agentic & Tool Use32.3#673
Reasoning23.1#1836
Math19.2#2885
Knowledge28.6#2456
Multimodal26.5#1204
Multilingual43.2#1851
Instruction Following68.8#1823
Long Context39.9#1671
Writing & Preference52.9#1647

Strengths and weaknesses

Categories where Claude 3.5 Sonnet places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude 3.5 Sonnet: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use32.3+2.0#67 of 154, top 44%
Coding39.0+0.2#165 of 340, top 49%
Reasoning23.1−0.5#183 of 350, top 53%

Weakest categories

Claude 3.5 Sonnet: weakest categories
CategoryScorevs medianRank
Multimodal26.5−12.0#120 of 128, top 94%
Math19.2−17.4#288 of 327, top 89%
Knowledge28.6−8.7#245 of 314, top 79%

Closest competitors

The models ranked just above and below Claude 3.5 Sonnet. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude 3.5 Sonnet
ModelRankScoreBlended $/MSpeed
Magistral Medium#22735.2$2.750Compare
Gemini 2.0 Flash (Feb 2025)#22835.1—92Compare
C4ai Aya Expanse 8b#22934.9——Compare
Qwen Max#23034.7$2.80—Compare
Qwen3 Coder Next#23234.3$0.29—Compare
Devstral Small 2505#23334.3$0.1588Compare
Qwen1.5-110B#23434.2——Compare
o1-mini#23534.0——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude 3.5 Sonnet Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot51.6%#22 of 44, top 50%Epoch AI
GSO4.6%#24 of 31, top 78%Epoch AI
WeirdML31%Epoch AI
WeirdML40%#77 of 119, top 65%Epoch AI
BigCodeBench Instruct44.6%BigCodeBench2024-10-22
BigCodeBench Instruct46.8%#11 of 64, top 18%BigCodeBench2024-06-20
LiveBench Coding67.1%#9 of 39, top 24%Epoch AI
LMArena Coding1342#176 of 294, top 60%LMArena2026-10-08
LMArena Coding1306LMArena2026-10-08
BigCodeBench Complete58.6%#8 of 66, top 13%BigCodeBench2024-06-20
BigCodeBench Complete57.5%BigCodeBench2024-10-22
CadEval48%#7 of 14, top 50%Epoch AI
HumanEval+81.7%#10 of 45, top 23%june 2024EvalPlus
MBPP+74.3%#6 of 38, top 16%june 2024EvalPlus

Agentic & Tool Use

Claude 3.5 Sonnet Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
TheAgentCompany24%#6 of 14, top 43%Epoch AI
Cybench17.5%#12 of 21, top 58%Epoch AI
BALROG32.6%#15 of 35, top 43%Epoch AI
METR Time Horizons45.2%#25 of 32, top 79%Epoch AI
METR Time Horizons40.2%Epoch AI

Reasoning

Claude 3.5 Sonnet Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench27.5%Epoch AI
SimpleBench41.4%#51 of 77, top 67%Epoch AI
EnigmaEval0.9%#32 of 38, top 85%Epoch AI
LiveBench Reasoning56.7%#14 of 39, top 36%Epoch AI
LMArena Hard Prompts1305#186 of 297, top 63%LMArena2026-10-08
LMArena Hard Prompts1275LMArena2026-10-08
DTBench67.8%Epoch AI
DTBench67.8%#98 of 151, top 65%Epoch AI
LiveBench Data Analysis55%#18 of 39, top 47%Epoch AI
Epoch Capabilities Index130Epoch AI2024-06-20
Epoch Capabilities Index133.55#130 of 213, top 62%Epoch AI2024-10-22
ForecastBench60.7#22 of 72, top 31%Epoch AI
ForecastBench59.4Epoch AI
LiveBench59%#13 of 39, top 34%Epoch AI

Math

Claude 3.5 Sonnet Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20258.5%#134 of 173, top 78%Epoch AI2025-02-25
OTIS Mock AIME 2024-20256.5%Epoch AI2025-02-25
Omni-MATH27.6%#43 of 57, top 76%HELM Capabilities
LiveBench Math52.3%#19 of 39, top 49%Epoch AI
LMArena Math1307#179 of 285, top 63%LMArena2026-10-08
LMArena Math1303LMArena2026-10-08
MATH Level 556.9%#40 of 79, top 51%Epoch AI2025-01-27
MATH Level 551.7%Epoch AI2025-01-27
FrontierMath (Feb 2025 set)2.1%#54 of 68, top 80%Epoch AI2025-03-06
FrontierMath (Feb 2025 set)1%Epoch AI2025-03-07
FrontierMath Tier 4 (v1)0%#46 of 55, top 84%Epoch AI2025-07-01
FrontierMath Tier 4 (v1)0%#46 of 55, top 84%Epoch AI2025-07-01

Knowledge

Claude 3.5 Sonnet Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond55.3%#122 of 186, top 66%Epoch AI2025-01-27
GPQA Diamond54%Epoch AI2025-01-27
Humanity's Last Exam4.1%#39 of 41, top 96%Epoch AI
MMLU-Pro77.7%#24 of 58, top 42%HELM Capabilities
Confabulations (lower is better)19.9%#32 of 51, top 63%Lech Mazur benchmarks
GPQA (HELM)56.5%#26 of 57, top 46%HELM Capabilities
LMArena Expert1265#189 of 273, top 70%LMArena2026-10-08
LMArena Expert1247LMArena2026-10-08
MMLU87.3%#2 of 81, top 3%Epoch AI
MMLU86.5%Epoch AI

Multimodal

Claude 3.5 Sonnet Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1125#104 of 122, top 86%LMArena2026-10-09
LMArena Vision1120LMArena2026-10-09
Video-MME60%#12 of 15, top 80%Epoch AI
Video-MME60%#12 of 15, top 80%Epoch AI
GeoBench62%#16 of 25, top 64%Epoch AI
VPCT33%#22 of 24, top 92%Epoch AI
VPCT33%#22 of 24, top 92%Epoch AI

Multilingual

Claude 3.5 Sonnet Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1283#185 of 297, top 63%LMArena2026-10-08
LMArena Non-English1270LMArena2026-10-08
LMArena Chinese1265LMArena2026-10-08
LMArena Chinese1272#198 of 285, top 70%LMArena2026-10-08
LMArena French1294LMArena2026-10-08
LMArena French1305#159 of 223, top 72%LMArena2026-10-08
LMArena German1297#151 of 231, top 66%LMArena2026-10-08
LMArena German1271LMArena2026-10-08
LMArena Japanese1234#145 of 211, top 69%LMArena2026-10-08
LMArena Japanese1231LMArena2026-10-08
LMArena Korean1200#162 of 213, top 77%LMArena2026-10-08
LMArena Korean1197LMArena2026-10-08
LMArena Russian1283LMArena2026-10-08
LMArena Russian1306#172 of 283, top 61%LMArena2026-10-08
LMArena Spanish1290#165 of 226, top 74%LMArena2026-10-08
LMArena Spanish1286LMArena2026-10-08

Instruction Following

Claude 3.5 Sonnet Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following69.3%#19 of 39, top 49%Epoch AI
IFEval85.5%#17 of 57, top 30%HELM Capabilities
LMArena Instruction Following1270LMArena2026-10-08
LMArena Instruction Following1297#182 of 298, top 62%LMArena2026-10-08

Long Context

Claude 3.5 Sonnet Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1275LMArena2026-10-08
LMArena Longer Query1311#180 of 291, top 62%LMArena2026-10-08

Writing & Preference

Claude 3.5 Sonnet Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1281LMArena2026-10-08
LMArena Text1298#193 of 297, top 65%LMArena2026-10-08
LMArena Creative Writing1292#172 of 295, top 59%LMArena2026-10-08
LMArena Creative Writing1238LMArena2026-10-08
Short-Story Creative Writing80.3%#14 of 39, top 36%Epoch AI
EQ-Bench Creative Writing1451#63 of 115, top 55%EQ-Bench
WildBench79.2%#33 of 57, top 58%HELM Capabilities
LMArena Multi-Turn1299LMArena2026-10-08
LMArena Multi-Turn1326#174 of 295, top 59%LMArena2026-10-08
LiveBench Language53.8%#7 of 39, top 18%Epoch AI

Compare Claude 3.5 Sonnet

Other Anthropic models

Frequently asked questions

How good is Claude 3.5 Sonnet?

Claude 3.5 Sonnet by Anthropic ranks 231st of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.6. Its strongest category is agentic & tool use, where it ranks 67th.

Is Claude 3.5 Sonnet open source?

No. Claude 3.5 Sonnet is proprietary and available only through Anthropic's API and partner platforms.

What are Claude 3.5 Sonnet's strengths and weaknesses?

Relative to other ranked models, Claude 3.5 Sonnet places best in agentic & tool use, coding, reasoning and lowest in multimodal, math, knowledge.

What is Claude 3.5 Sonnet best at?

Its best category is agentic & tool use, where it ranks 67th on Noometry.