Model comparison

Claude 3.5 Sonnet vs Claude Fable 5.1

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 34.6 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 0 categories and Claude Fable 5.1 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Fable 5.1 leads 89.6 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 100% for Claude Fable 5.1.

Side by side

Claude 3.5 Sonnet and Claude Fable 5.1 specifications
Claude 3.5 SonnetClaude Fable 5.1
ProviderAnthropicAnthropic
Noometry Index34.669.0
Released2024-06-202026-09-01
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$10
Output $ / M tokens—$50
Results tracked6052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude 3.5 Sonnet: 39.0 (#165), Claude Fable 5.1: 74.7 (#1)

Coding benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
GSO4.6%88.2%
WeirdML40%92.9%
LMArena Coding13421528
FrontierCode—50.9%
Aider Polyglot51.6%—
CursorBench—51.8%
LMArena WebDev—1744
FrontierSWE—56.3%
SciCode—63.1%
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
MirrorCode—73.3%
BigCodeBench Complete58.6%—
CadEval48%—
ALE-Bench—2,143
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Claude Fable 5.1 leads

Claude 3.5 Sonnet: 32.3 (#67), Claude Fable 5.1: 50.7 (#5)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
APEX-Agents—68.6%
Remote Labor Index—17.9%
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
GDP.pdf—29.6%
METR Time Horizons45.2%—
Vending-Bench 2—5,422

Reasoning Claude Fable 5.1 leads

Claude 3.5 Sonnet: 23.1 (#183), Claude Fable 5.1: 76.7 (#7)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
LMArena Hard Prompts13051526
DTBench67.8%97.6%
Epoch Capabilities Index133.55164.7
ARC-AGI-2—90%
SimpleBench41.4%—
NYT Connections (extended)—90%
ARC-AGI-1—97.5%
CritPt—31.1%
Chess Puzzles—47%
EnigmaEval0.9%—
EBR-Bench—57.1%
LiveBench Reasoning56.7%—
Mystery Game Puzzles—58%
LiveBench Data Analysis55%—
LMCA—65.5%
ForecastBench60.7—
LiveBench59%—

Math Claude Fable 5.1 leads

Claude 3.5 Sonnet: 19.2 (#288), Claude Fable 5.1: 89.6 (#4)

Math benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
OTIS Mock AIME 2024-20258.5%100%
LMArena Math13071525
FrontierMath (Tiers 1-3)—90.2%
FrontierMath Tier 4—87.8%
ProofBench—100%
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Erdős—0%
FrontierMath Tier 4 (v1)0%—

Knowledge Claude Fable 5.1 leads

Claude 3.5 Sonnet: 28.6 (#245), Claude Fable 5.1: 69.6 (#6)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
Humanity's Last Exam4.1%46.5%
LMArena Expert12651535
GPQA Diamond55.3%—
SimpleQA Verified—70.8%
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Claude Fable 5.1 leads

Claude 3.5 Sonnet: 26.5 (#120), Claude Fable 5.1: 53.9 (#4)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
LMArena Vision11251318
Video-MME60%—
GeoBench62%—
VPCT33%—
Blueprint-Bench 2—41.9%
Furniture Assembly—70%
LMArena Document—1513

Multilingual Claude Fable 5.1 leads

Claude 3.5 Sonnet: 43.2 (#185), Claude Fable 5.1: 59.1 (#3)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
LMArena Non-English12831507
LMArena Chinese12721586
LMArena French13051525
LMArena German12971500
LMArena Japanese12341543
LMArena Korean12001534
LMArena Russian13061521
LMArena Spanish12901516

Instruction Following Claude Fable 5.1 leads

Claude 3.5 Sonnet: 68.8 (#182), Claude Fable 5.1: 79.2 (#6)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
LMArena Instruction Following12971517
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Claude Fable 5.1 leads

Claude 3.5 Sonnet: 39.9 (#167), Claude Fable 5.1: 46.7 (#20)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
LMArena Longer Query13111522

Writing & Preference Claude Fable 5.1 leads

Claude 3.5 Sonnet: 52.9 (#164), Claude Fable 5.1: 79.2 (#2)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetClaude Fable 5.1
LMArena Text12981510
LMArena Creative Writing12921507
EQ-Bench Creative Writing14512162
LMArena Multi-Turn13261492
Short-Story Creative Writing80.3%—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Claude Fable 5.1?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Claude Fable 5.1 better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Claude Fable 5.1 share?

25 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Claude Fable 5.1 has 52.

Related comparisons

Go deeper