Model comparison

Claude 3.7 Sonnet vs Claude Fable 5.1

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 39.5 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 1 category and Claude Fable 5.1 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Fable 5.1 leads 76.7 to 18.6.
  • The biggest single-benchmark swing is ARC-AGI-2: 0.9% for Claude 3.7 Sonnet and 90% for Claude Fable 5.1.

Side by side

Claude 3.7 Sonnet and Claude Fable 5.1 specifications
Claude 3.7 SonnetClaude Fable 5.1
ProviderAnthropicAnthropic
Noometry Index39.569.0
Released2025-02-242026-09-01
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$10
Output $ / M tokens—$50
Results tracked5852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude 3.7 Sonnet: 40.6 (#136), Claude Fable 5.1: 74.7 (#1)

Coding benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
GSO3.8%88.2%
LMArena Coding13611528
SWE-bench Verified61%—
FrontierCode—50.9%
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
CursorBench—51.8%
LMArena WebDev—1744
FrontierSWE—56.3%
SciCode—63.1%
WeirdML—92.9%
LiveBench Coding74.5%—
MirrorCode—73.3%
CadEval54%—
ALE-Bench—2,143

Agentic & Tool Use Claude Fable 5.1 leads

Claude 3.7 Sonnet: 34.1 (#50), Claude Fable 5.1: 50.7 (#5)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
APEX-Agents—68.6%
Remote Labor Index—17.9%
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
GDP.pdf—29.6%
METR Time Horizons60%—
Vending-Bench 2—5,422

Reasoning Claude Fable 5.1 leads

Claude 3.7 Sonnet: 18.6 (#277), Claude Fable 5.1: 76.7 (#7)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
ARC-AGI-20.9%90%
ARC-AGI-128.6%97.5%
LMArena Hard Prompts13331526
Epoch Capabilities Index141.16164.7
SimpleBench46.4%—
NYT Connections (extended)—90%
CritPt—31.1%
Chess Puzzles—47%
EnigmaEval4.2%—
EBR-Bench—57.1%
LiveBench Reasoning87.8%—
Mystery Game Puzzles—58%
DTBench—97.6%
LiveBench Data Analysis74%—
LMCA—65.5%
ForecastBench61.8—
LiveBench76.1%—

Math Claude Fable 5.1 leads

Claude 3.7 Sonnet: 37.5 (#153), Claude Fable 5.1: 89.6 (#4)

Math benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
OTIS Mock AIME 2024-202557.8%100%
LMArena Math13371525
FrontierMath (Tiers 1-3)—90.2%
FrontierMath Tier 4—87.8%
ProofBench—100%
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—
FrontierMath Erdős—0%

Knowledge Claude Fable 5.1 leads

Claude 3.7 Sonnet: 39.8 (#130), Claude Fable 5.1: 69.6 (#6)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
Humanity's Last Exam8%46.5%
LMArena Expert13211535
GPQA Diamond79.7%—
SimpleQA Verified—70.8%
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—

Multimodal Claude Fable 5.1 leads

Claude 3.7 Sonnet: 33.7 (#95), Claude Fable 5.1: 53.9 (#4)

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
LMArena Vision11691318
GeoBench68%—
VPCT39%—
Blueprint-Bench 2—41.9%
Furniture Assembly—70%
LMArena Document—1513
SpatialViz-Bench33.9%—

Multilingual Claude Fable 5.1 leads

Claude 3.7 Sonnet: 44.1 (#179), Claude Fable 5.1: 59.1 (#3)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
LMArena Non-English12961507
LMArena Chinese12991586
LMArena French13031525
LMArena German13011500
LMArena Japanese12671543
LMArena Korean12491534
LMArena Russian13111521
LMArena Spanish12981516

Instruction Following Claude Fable 5.1 leads

Claude 3.7 Sonnet: 72.9 (#125), Claude Fable 5.1: 79.2 (#6)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
LMArena Instruction Following13521517
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Claude Fable 5.1: 46.7 (#20)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
LMArena Longer Query13731522
Fiction.LiveBench83.3%—

Writing & Preference Claude Fable 5.1 leads

Claude 3.7 Sonnet: 54.4 (#150), Claude Fable 5.1: 79.2 (#2)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetClaude Fable 5.1
LMArena Text13141510
LMArena Creative Writing13321507
EQ-Bench Creative Writing14122162
LMArena Multi-Turn13391492
Short-Story Creative Writing81.1%—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Claude Fable 5.1?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or Claude Fable 5.1 better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Claude Fable 5.1 share?

25 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Claude Fable 5.1 has 52.

Related comparisons

Go deeper