Model comparison

Llama 3.1 Nemotron 70b Instruct vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 37.6 on the Noometry Index.

Last verified . 12 shared benchmarks.

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Llama 3.1 Nemotron 70b Instruct scores higher in 2 categories and Trinity Large Thinking in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.1 Nemotron 70b Instruct leads 25.0 to 16.9.

Side by side

Llama 3.1 Nemotron 70b Instruct and Trinity Large Thinking specifications
Llama 3.1 Nemotron 70b InstructTrinity Large Thinking
ProviderNVIDIAArcee AI
Noometry Index37.638.6
Released2024-12-182026-04-01
WeightsOpenOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked1424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 35.9 (#216), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Coding12721381
LMArena WebDev—1238
SciCode—36.1%
BigCodeBench Instruct38.7%—
BigCodeBench Complete48.2%—

Reasoning Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 25.0 (#152), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Hard Prompts12661350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Llama 3.1 Nemotron 70b Instruct: 35.5 (#182), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Math12711366

Knowledge Trinity Large Thinking leads

Llama 3.1 Nemotron 70b Instruct: 34.1 (#199), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Expert12421360
Vectara Hallucination Rate—6.9%

Multilingual Trinity Large Thinking leads

Llama 3.1 Nemotron 70b Instruct: 40.5 (#217), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Non-English12451325
LMArena Chinese12631373
LMArena Russian12271337
LMArena French—1374
LMArena German—1356
LMArena Japanese—1311
LMArena Korean—1306
LMArena Spanish—1357

Instruction Following Trinity Large Thinking leads

Llama 3.1 Nemotron 70b Instruct: 65.9 (#213), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Instruction Following12521334

Long Context Trinity Large Thinking leads

Llama 3.1 Nemotron 70b Instruct: 37.6 (#215), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Longer Query12381355

Writing & Preference Trinity Large Thinking leads

Llama 3.1 Nemotron 70b Instruct: 48.4 (#203), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructTrinity Large Thinking
LMArena Text12831340
LMArena Creative Writing12691320
LMArena Multi-Turn12751342

Frequently asked questions

Is Llama 3.1 Nemotron 70b Instruct better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 37.6 on the Noometry Index.

Is Llama 3.1 Nemotron 70b Instruct or Trinity Large Thinking better for coding?

Llama 3.1 Nemotron 70b Instruct scores higher on coding benchmarks: 35.9 versus 34.1 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron 70b Instruct and Trinity Large Thinking share?

12 benchmarks have published results for both models. Llama 3.1 Nemotron 70b Instruct has 14 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper