Model comparison

Claude Fable 5.1 vs GPT-6 Astra

GPT-6 Astra is the stronger model overall, scoring 70.8 to 69.0 on the Noometry Index.

Last verified . 51 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

GPT-6 Astra OpenAI

70.8

Rank #1 Confirmed

Summary

  • They share 51 benchmarks with published results for both. Claude Fable 5.1 scores higher in 5 categories and GPT-6 Astra in 5 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-6 Astra leads 85.1 to 76.7.
  • The biggest single-benchmark swing is MirrorCode: 73.3% for Claude Fable 5.1 and 46.7% for GPT-6 Astra.
  • Both cost about the same: $10 input and $50 output per million tokens.
  • GPT-6 Astra accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Fable 5.1 and GPT-6 Astra specifications
Claude Fable 5.1GPT-6 Astra
ProviderAnthropicOpenAI
Noometry Index69.070.8
Released2026-09-012026-09-03
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K128K
Input $ / M tokens$10$10
Output $ / M tokens$50$50
Results tracked5256

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), GPT-6 Astra: 73.7 (#2)

Coding benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
FrontierCode50.9%53.3%
LMArena WebDev17441786
FrontierSWE56.3%65.5%
SciCode63.1%56.5%
GSO88.2%79.4%
WeirdML92.9%93.6%
LMArena Coding15281487
MirrorCode73.3%46.7%
ALE-Bench2,1432,951
DeepSWE—74.1%
CursorBench51.8%—

Agentic & Tool Use GPT-6 Astra leads

Claude Fable 5.1: 50.7 (#5), GPT-6 Astra: 52.9 (#3)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
APEX-Agents68.6%64.7%
Remote Labor Index17.9%20.8%
GDP.pdf29.6%34.2%
Vending-Bench 25,42215,515
BALROG—68.3%

Reasoning GPT-6 Astra leads

Claude Fable 5.1: 76.7 (#7), GPT-6 Astra: 85.1 (#1)

Reasoning benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
ARC-AGI-290%95%
NYT Connections (extended)90%98.1%
ARC-AGI-197.5%98.5%
CritPt31.1%31.7%
Chess Puzzles47%72%
EBR-Bench57.1%76.2%
LMArena Hard Prompts15261462
Mystery Game Puzzles58%84%
DTBench97.6%97.3%
LMCA65.5%64.4%
Epoch Capabilities Index164.7166.45
Bench to the Future 3—0.14

Math GPT-6 Astra leads

Claude Fable 5.1: 89.6 (#4), GPT-6 Astra: 93.5 (#2)

Math benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
FrontierMath (Tiers 1-3)90.2%93.7%
FrontierMath Tier 487.8%97.6%
OTIS Mock AIME 2024-2025100%100%
ProofBench100%99%
LMArena Math15251465
FrontierMath Erdős0%2.9%

Knowledge GPT-6 Astra leads

Claude Fable 5.1: 69.6 (#6), GPT-6 Astra: 75.3 (#1)

Knowledge benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
Humanity's Last Exam46.5%54.8%
SimpleQA Verified70.8%75.6%
LMArena Expert15351483
GPQA Diamond—95.8%
Vectara Hallucination Rate—8.7%

Multimodal GPT-6 Astra leads

Claude Fable 5.1: 53.9 (#4), GPT-6 Astra: 55.0 (#3)

Multimodal benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
LMArena Vision13181281
Blueprint-Bench 241.9%49.7%
Furniture Assembly70%80%
LMArena Document15131468

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), GPT-6 Astra: 53.7 (#61)

Multilingual benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
LMArena Non-English15071430
LMArena Chinese15861484
LMArena French15251456
LMArena German15001440
LMArena Japanese15431379
LMArena Korean15341426
LMArena Russian15211436
LMArena Spanish15161407

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), GPT-6 Astra: 76.3 (#44)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
LMArena Instruction Following15171450

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), GPT-6 Astra: 44.5 (#62)

Long Context benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
LMArena Longer Query15221456

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), GPT-6 Astra: 75.3 (#7)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1GPT-6 Astra
LMArena Text15101441
LMArena Creative Writing15071418
EQ-Bench Creative Writing21622173
LMArena Multi-Turn14921448

Frequently asked questions

Is Claude Fable 5.1 better than GPT-6 Astra?

GPT-6 Astra is the stronger model overall, scoring 70.8 to 69.0 on the Noometry Index.

Which is cheaper, Claude Fable 5.1 or GPT-6 Astra?

GPT-6 Astra is cheaper. It lists at $10 per million input tokens and $50 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or GPT-6 Astra better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 73.7 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Astra does, with 1.05M tokens against 1M.

How many benchmarks do Claude Fable 5.1 and GPT-6 Astra share?

51 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and GPT-6 Astra has 56.

Related comparisons

Go deeper