Anthropic, proprietary

Claude 3 Sonnet

Claude 3 Sonnet by Anthropic ranks 319th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is multimodal, where it ranks 125th.

Last verified

Specifications

Noometry rank
#319 of 354
Index score
29.0
Evidence
Confirmed 30 results
Provider
Anthropic
Released
February 29, 2024
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Claude 3 Sonnet category scores
  1. Coding 29.6
  2. Reasoning 20.5
  3. Math 10.7
  4. Knowledge 21.1
  5. Multimodal 25.2
  6. Multilingual 37.8
  7. Instruction Following 62.8
  8. Long Context 36.7
  9. Writing & Preference 42.1
Claude 3 Sonnet category ranks
CategoryScoreRankResults
Coding29.6#3024
Reasoning20.5#2372
Math10.7#3103
Knowledge21.1#2762
Multimodal25.2#1251
Multilingual37.8#2341
Instruction Following62.8#2351
Long Context36.7#2281
Writing & Preference42.1#2383

Strengths and weaknesses

Categories where Claude 3 Sonnet places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Claude 3 Sonnet: strongest categories
CategoryScorevs medianRank
Reasoning20.5−3.1#237 of 350, top 68%
Writing & Preference42.1−11.7#238 of 312, top 77%
Long Context36.7−4.2#228 of 296, top 78%

Weakest categories

Claude 3 Sonnet: weakest categories
CategoryScorevs medianRank
Multimodal25.2−13.3#125 of 128, top 98%
Math10.7−25.9#310 of 327, top 95%
Coding29.6−9.1#302 of 340, top 89%

Closest competitors

The models ranked just above and below Claude 3 Sonnet. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Claude 3 Sonnet
ModelRankScoreBlended $/MSpeed
Claude 3.5 Haiku#31529.2——Compare
GPT-4#31629.1$37.50—Compare
Llama 2-7B#31729.1——Compare
Granite 4.0 Micro#31829.0$0.0408—Compare
Qwen2.5 7B Instruct#32029.0$0.31—Compare
Llama 3.2 3B#32128.9$0.12—Compare
Qwen1.5 4b Chat#32228.8——Compare
Llama 3-70B#32328.8—104Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Claude 3 Sonnet Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML10.2%#113 of 119, top 95%Epoch AI
BigCodeBench Instruct42.7%#28 of 64, top 44%BigCodeBench2024-02-29
LMArena Coding1223#233 of 294, top 80%LMArena2026-10-08
BigCodeBench Complete53.8%#24 of 66, top 37%BigCodeBench2024-02-29
HumanEval+64%#24 of 45, top 54%mar 2024EvalPlus
MBPP+69.3%#15 of 38, top 40%mar 2024EvalPlus

Reasoning

Claude 3 Sonnet Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1197#238 of 297, top 81%LMArena2026-10-08
DTBench53.6%#127 of 151, top 85%Epoch AI
Epoch Capabilities Index120.7#166 of 213, top 78%Epoch AI2024-02-29
WinoGrande75.1%#24 of 43, top 56%Epoch AI

Math

Claude 3 Sonnet Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20252.5%#158 of 173, top 92%Epoch AI2025-02-25
LMArena Math1213#227 of 285, top 80%LMArena2026-10-08
MATH Level 518.2%#64 of 79, top 82%Epoch AI2025-01-27

Knowledge

Claude 3 Sonnet Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond40.6%#149 of 186, top 81%Epoch AI2025-01-27
LMArena Expert1173#228 of 273, top 84%LMArena2026-10-08
MMLU75.9%#31 of 81, top 39%Epoch AI

Multimodal

Claude 3 Sonnet Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision984#119 of 122, top 98%LMArena2026-10-09

Multilingual

Claude 3 Sonnet Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1205#234 of 297, top 79%LMArena2026-10-08
LMArena Chinese1189#236 of 285, top 83%LMArena2026-10-08
LMArena French1229#189 of 223, top 85%LMArena2026-10-08
LMArena German1204#191 of 231, top 83%LMArena2026-10-08
LMArena Japanese1131#180 of 211, top 86%LMArena2026-10-08
LMArena Korean1128#186 of 213, top 88%LMArena2026-10-08
LMArena Russian1227#226 of 283, top 80%LMArena2026-10-08
LMArena Spanish1204#193 of 226, top 86%LMArena2026-10-08

Instruction Following

Claude 3 Sonnet Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1199#236 of 298, top 80%LMArena2026-10-08

Long Context

Claude 3 Sonnet Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1211#238 of 291, top 82%LMArena2026-10-08

Writing & Preference

Claude 3 Sonnet Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1218#238 of 297, top 81%LMArena2026-10-08
LMArena Creative Writing1186#238 of 295, top 81%LMArena2026-10-08
LMArena Multi-Turn1227#227 of 295, top 77%LMArena2026-10-08

Compare Claude 3 Sonnet

Other Anthropic models

Frequently asked questions

How good is Claude 3 Sonnet?

Claude 3 Sonnet by Anthropic ranks 319th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.0. Its strongest category is multimodal, where it ranks 125th.

Is Claude 3 Sonnet open source?

No. Claude 3 Sonnet is proprietary and available only through Anthropic's API and partner platforms.

What are Claude 3 Sonnet's strengths and weaknesses?

Relative to other ranked models, Claude 3 Sonnet places best in reasoning, writing & preference, long context and lowest in multimodal, math, coding.

What is Claude 3 Sonnet best at?

Its best category is multimodal, where it ranks 125th on Noometry.