Model comparison

Trinity Large Thinking vs Tulu 3 (Tülu 3) 70B

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 33.0 on the Noometry Index.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Trinity Large Thinking scores higher in 6 categories and Tulu 3 (Tülu 3) 70B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Trinity Large Thinking leads 37.6 to 14.2.

Side by side

Trinity Large Thinking and Tulu 3 (Tülu 3) 70B specifications
Trinity Large ThinkingTulu 3 (Tülu 3) 70B
ProviderArcee AIAllen Institute for AI (Ai2)
Noometry Index38.633.0
Released2026-04-012024-11-21
WeightsOpenOpen
Context window262K—
Max output80K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.80—
Results tracked2414

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Tulu 3 (Tülu 3) 70B leads

Trinity Large Thinking: 34.1 (#244), Tulu 3 (Tülu 3) 70B: 36.0 (#214)

Coding benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Coding13811235
LMArena WebDev1238—
SciCode36.1%—

Reasoning Tulu 3 (Tülu 3) 70B leads

Trinity Large Thinking: 16.9 (#298), Tulu 3 (Tülu 3) 70B: 23.9 (#169)

Reasoning benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Hard Prompts13501220
NYT Connections (extended)16.5%—
CritPt0.9%—
Thematic Generalization41.6%—
Surface Evolver Bench15.6%—

Math Trinity Large Thinking leads

Trinity Large Thinking: 37.6 (#149), Tulu 3 (Tülu 3) 70B: 14.2 (#303)

Math benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Math13661242
OTIS Mock AIME 2024-2025—4.4%
MATH Level 5—42.7%

Knowledge Trinity Large Thinking leads

Trinity Large Thinking: 40.9 (#113), Tulu 3 (Tülu 3) 70B: 25.0 (#264)

Knowledge benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
GPQA Diamond—46.3%
Vectara Hallucination Rate6.9%—
LMArena Expert1360—

Multilingual Trinity Large Thinking leads

Trinity Large Thinking: 46.2 (#160), Tulu 3 (Tülu 3) 70B: 39.9 (#222)

Multilingual benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Non-English13251236
LMArena Chinese13731249
LMArena Russian13371246
LMArena French1374—
LMArena German1356—
LMArena Japanese1311—
LMArena Korean1306—
LMArena Spanish1357—

Instruction Following Trinity Large Thinking leads

Trinity Large Thinking: 70.5 (#162), Tulu 3 (Tülu 3) 70B: 64.8 (#227)

Instruction Following benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Instruction Following13341233

Long Context Trinity Large Thinking leads

Trinity Large Thinking: 41.3 (#144), Tulu 3 (Tülu 3) 70B: 37.1 (#222)

Long Context benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Longer Query13551224

Writing & Preference Trinity Large Thinking leads

Trinity Large Thinking: 53.8 (#158), Tulu 3 (Tülu 3) 70B: 45.6 (#223)

Writing & Preference benchmarks
BenchmarkTrinity Large ThinkingTulu 3 (Tülu 3) 70B
LMArena Text13401256
LMArena Creative Writing13201231
LMArena Multi-Turn13421252

Frequently asked questions

Is Trinity Large Thinking better than Tulu 3 (Tülu 3) 70B?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 33.0 on the Noometry Index.

Is Trinity Large Thinking or Tulu 3 (Tülu 3) 70B better for coding?

Tulu 3 (Tülu 3) 70B scores higher on coding benchmarks: 36.0 versus 34.1 in the Noometry coding category.

How many benchmarks do Trinity Large Thinking and Tulu 3 (Tülu 3) 70B share?

11 benchmarks have published results for both models. Trinity Large Thinking has 24 scored results on Noometry and Tulu 3 (Tülu 3) 70B has 14.

Related comparisons

Go deeper