Model comparison

Claude Fable 5 vs Trinity Large Thinking

Claude Fable 5 is the stronger model overall, scoring 66.8 to 38.6 on the Noometry Index. Trinity Large Thinking costs 52× less per token, which makes it the better buy when Claude Fable 5's lead doesn't matter for your workload.

Last verified . 22 shared benchmarks.

Claude Fable 5 Anthropic

66.8

Rank #5 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude Fable 5 scores higher in 8 categories and Trinity Large Thinking in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Fable 5 leads 76.8 to 16.9.
  • The biggest single-benchmark swing is Surface Evolver Bench: 95% for Claude Fable 5 and 15.6% for Trinity Large Thinking.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $10 / $50 for Claude Fable 5.
  • Claude Fable 5 accepts more context: 1M tokens versus 262K.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Claude Fable 5 and Trinity Large Thinking specifications
Claude Fable 5Trinity Large Thinking
ProviderAnthropicArcee AI
Noometry Index66.838.6
Released2026-06-072026-04-01
WeightsProprietaryOpen
Context window1M262K
Max output128K80K
Input $ / M tokens$10$0.25
Output $ / M tokens$50$0.80
Results tracked6224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5 leads

Claude Fable 5: 70.6 (#4), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena WebDev16251238
SciCode61%36.1%
LMArena Coding15191381
DeepSWE69.9%—
FrontierCode53.5%—
FrontierSWE47%—
GSO78.4%—
WeirdML91.9%—
MirrorCode63.9%—
ALE-Bench2,041—

Agentic & Tool Use Not comparable

Claude Fable 5: 54.0 (#2), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
APEX-Agents63.6%—
Remote Labor Index16.1%—
τ²-bench Banking39.7%—
PostTrainBench41.8%—
GBAEval74.5%—
GDP.pdf30%—
LMArena Search1230—
Vending-Bench 25,680—

Reasoning Claude Fable 5 leads

Claude Fable 5: 76.8 (#6), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
NYT Connections (extended)92.7%16.5%
CritPt28.6%0.9%
LMArena Hard Prompts15081350
Surface Evolver Bench95%15.6%
ARC-AGI-289.2%—
SimpleBench81.9%—
Kagi LLM Benchmark91.4%—
ARC-AGI-198.5%—
Chess Puzzles41%—
EnigmaEval39.3%—
Thematic Generalization—41.6%
EBR-Bench39.5%—
Mystery Game Puzzles52%—
DTBench98.4%—
LMCA61.1%—
Bench to the Future 30.13—
Epoch Capabilities Index162.06—

Math Claude Fable 5 leads

Claude Fable 5: 88.5 (#5), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Math15191366
FrontierMath (Tiers 1-3)87%—
FrontierMath Tier 490.2%—
OTIS Mock AIME 2024-2025100%—
ProofBench95%—
FrontierMath Erdős0%—

Knowledge Claude Fable 5 leads

Claude Fable 5: 62.2 (#25), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Expert15341360
GPQA Diamond85.9%—
SimpleQA Verified70.7%—
Vectara Hallucination Rate—6.9%

Multimodal Not comparable

Claude Fable 5: 45.3 (#17), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Vision1324—
Blueprint-Bench 238.6%—
Furniture Assembly35.8%—
LMArena Document1496—

Multilingual Claude Fable 5 leads

Claude Fable 5: 57.3 (#9), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Non-English14811325
LMArena Chinese15431373
LMArena French15051374
LMArena German14861356
LMArena Japanese15061311
LMArena Korean14881306
LMArena Russian15041337
LMArena Spanish14981357

Instruction Following Claude Fable 5 leads

Claude Fable 5: 78.6 (#8), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Instruction Following15021334

Long Context Claude Fable 5 leads

Claude Fable 5: 46.3 (#23), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Longer Query15091355

Writing & Preference Claude Fable 5 leads

Claude Fable 5: 75.9 (#5), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkClaude Fable 5Trinity Large Thinking
LMArena Text14911340
LMArena Creative Writing14941320
LMArena Multi-Turn15041342
EQ-Bench Creative Writing1943—
EQ-Bench 41340—

Frequently asked questions

Is Claude Fable 5 better than Trinity Large Thinking?

Claude Fable 5 is the stronger model overall, scoring 66.8 to 38.6 on the Noometry Index. Trinity Large Thinking costs 52× less per token, which makes it the better buy when Claude Fable 5's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5 or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Claude Fable 5 lists at $10 and $50.

Is Claude Fable 5 or Trinity Large Thinking better for coding?

Claude Fable 5 scores higher on coding benchmarks: 70.6 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Claude Fable 5 does, with 1M tokens against 262K.

How many benchmarks do Claude Fable 5 and Trinity Large Thinking share?

22 benchmarks have published results for both models. Claude Fable 5 has 62 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper