Model comparison

Claude 3.7 Sonnet vs Trinity Large Thinking

Claude 3.7 Sonnet and Trinity Large Thinking score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 5 categories and Trinity Large Thinking in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Claude 3.7 Sonnet leads 50.3 to 41.3.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Claude 3.7 Sonnet and Trinity Large Thinking specifications
Claude 3.7 SonnetTrinity Large Thinking
ProviderAnthropicArcee AI
Noometry Index39.538.6
Released2025-02-242026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked5824

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Coding13611381
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
Aider Polyglot64.9%—
LMArena WebDev—1238
SciCode—36.1%
GSO3.8%—
LiveBench Coding74.5%—
CadEval54%—

Agentic & Tool Use Not comparable

Claude 3.7 Sonnet: 34.1 (#50), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
METR Time Horizons60%—

Reasoning Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 18.6 (#277), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Hard Prompts13331350
ARC-AGI-20.9%—
SimpleBench46.4%—
NYT Connections (extended)—16.5%
ARC-AGI-128.6%—
CritPt—0.9%
EnigmaEval4.2%—
Thematic Generalization—41.6%
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
Surface Evolver Bench—15.6%
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Too close to call

Claude 3.7 Sonnet: 37.5 (#153), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Math13371366
OTIS Mock AIME 2024-202557.8%—
Omni-MATH33%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Trinity Large Thinking leads

Claude 3.7 Sonnet: 39.8 (#130), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Expert13211360
GPQA Diamond79.7%—
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
Vectara Hallucination Rate—6.9%
GPQA (HELM)60.8%—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Trinity Large Thinking leads

Claude 3.7 Sonnet: 44.1 (#179), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Non-English12961325
LMArena Chinese12991373
LMArena French13031374
LMArena German13011356
LMArena Japanese12671311
LMArena Korean12491306
LMArena Russian13111337
LMArena Spanish12981357

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Instruction Following13521334
LiveBench Instruction Following81.3%—
IFEval83.4%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Longer Query13731355
Fiction.LiveBench83.3%—

Writing & Preference Too close to call

Claude 3.7 Sonnet: 54.4 (#150), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetTrinity Large Thinking
LMArena Text13141340
LMArena Creative Writing13321320
LMArena Multi-Turn13391342
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Trinity Large Thinking?

Claude 3.7 Sonnet and Trinity Large Thinking score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Is Claude 3.7 Sonnet or Trinity Large Thinking better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 34.1 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Trinity Large Thinking share?

17 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper