Model comparison

Olmo 7b Instruct vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

Summary

  • They share 10 benchmarks with published results for both. Olmo 7b Instruct scores higher in 1 category and Trinity Large Thinking in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Trinity Large Thinking leads 53.8 to 25.8.

Side by side

Olmo 7b Instruct and Trinity Large Thinking specifications
Olmo 7b InstructTrinity Large Thinking
ProviderAllen Institute for AI (Ai2)Arcee AI
Noometry Index30.338.6
Released—2026-04-01
WeightsOpenOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked1024

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Trinity Large Thinking leads

Olmo 7b Instruct: 29.6 (#303), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Coding10161381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Olmo 7b Instruct leads

Olmo 7b Instruct: 18.8 (#274), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Hard Prompts9931350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Olmo 7b Instruct: 30.2 (#237), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Math10181366

Knowledge Not comparable

Olmo 7b Instruct: —, Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
Vectara Hallucination Rate—6.9%
LMArena Expert—1360

Multilingual Trinity Large Thinking leads

Olmo 7b Instruct: 24.0 (#291), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Non-English9771325
LMArena Chinese10141373
LMArena Russian9471337
LMArena French—1374
LMArena German—1356
LMArena Japanese—1311
LMArena Korean—1306
LMArena Spanish—1357

Instruction Following Trinity Large Thinking leads

Olmo 7b Instruct: 49.0 (#301), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Instruction Following9781334

Long Context Not comparable

Olmo 7b Instruct: —, Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Longer Query—1355

Writing & Preference Trinity Large Thinking leads

Olmo 7b Instruct: 25.8 (#303), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkOlmo 7b InstructTrinity Large Thinking
LMArena Text10321340
LMArena Creative Writing9901320
LMArena Multi-Turn10071342

Frequently asked questions

Is Olmo 7b Instruct better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 30.3 on the Noometry Index.

Is Olmo 7b Instruct or Trinity Large Thinking better for coding?

Trinity Large Thinking scores higher on coding benchmarks: 34.1 versus 29.6 in the Noometry coding category.

How many benchmarks do Olmo 7b Instruct and Trinity Large Thinking share?

10 benchmarks have published results for both models. Olmo 7b Instruct has 10 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper