Model comparison

Llama2 70b Steerlm Chat vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 31.8 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 1 category and Trinity Large Thinking in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Trinity Large Thinking leads 53.8 to 31.6.

Side by side

Llama2 70b Steerlm Chat and Trinity Large Thinking specifications
Llama2 70b Steerlm ChatTrinity Large Thinking
ProviderNVIDIAArcee AI
Noometry Index31.838.6
Released—2026-04-01
WeightsOpenOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked924

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Trinity Large Thinking leads

Llama2 70b Steerlm Chat: 29.9 (#300), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Coding10251381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 20.0 (#246), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Hard Prompts10471350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Llama2 70b Steerlm Chat: 31.3 (#226), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Math10721366

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
Vectara Hallucination Rate—6.9%
LMArena Expert—1360

Multilingual Trinity Large Thinking leads

Llama2 70b Steerlm Chat: 28.8 (#270), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Non-English10631325
LMArena Chinese—1373
LMArena French—1374
LMArena German—1356
LMArena Japanese—1311
LMArena Korean—1306
LMArena Russian—1337
LMArena Spanish—1357

Instruction Following Trinity Large Thinking leads

Llama2 70b Steerlm Chat: 54.2 (#279), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Instruction Following10601334

Long Context Trinity Large Thinking leads

Llama2 70b Steerlm Chat: 30.4 (#288), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Longer Query9981355

Writing & Preference Trinity Large Thinking leads

Llama2 70b Steerlm Chat: 31.6 (#283), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatTrinity Large Thinking
LMArena Text10981340
LMArena Creative Writing10911320
LMArena Multi-Turn10581342

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 31.8 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Trinity Large Thinking better for coding?

Trinity Large Thinking scores higher on coding benchmarks: 34.1 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Trinity Large Thinking share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper