Model comparison

Olmo 3.1 32b Instruct vs Trinity Large Thinking

Olmo 3.1 32b Instruct and Trinity Large Thinking score almost the same on the Noometry Index (39.4 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 2 categories and Trinity Large Thinking in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 16.9.

Side by side

Olmo 3.1 32b Instruct and Trinity Large Thinking specifications
Olmo 3.1 32b InstructTrinity Large Thinking
ProviderAllen Institute for AI (Ai2)Arcee AI
Noometry Index39.438.6
Released—2026-04-01
WeightsOpenOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked1624

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Coding13471381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Hard Prompts13221350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Olmo 3.1 32b Instruct: 36.3 (#167), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Math13051366

Knowledge Trinity Large Thinking leads

Olmo 3.1 32b Instruct: 36.1 (#175), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Expert13081360
Vectara Hallucination Rate—6.9%

Multilingual Trinity Large Thinking leads

Olmo 3.1 32b Instruct: 42.6 (#191), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Non-English12751325
LMArena Chinese13041373
LMArena French13281374
LMArena German12821356
LMArena Korean12061306
LMArena Russian12681337
LMArena Spanish13361357
LMArena Japanese—1311

Instruction Following Trinity Large Thinking leads

Olmo 3.1 32b Instruct: 68.6 (#187), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Instruction Following12991334

Long Context Trinity Large Thinking leads

Olmo 3.1 32b Instruct: 39.9 (#166), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Longer Query13121355

Writing & Preference Trinity Large Thinking leads

Olmo 3.1 32b Instruct: 50.2 (#185), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructTrinity Large Thinking
LMArena Text13111340
LMArena Creative Writing12641320
LMArena Multi-Turn13091342

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Trinity Large Thinking?

Olmo 3.1 32b Instruct and Trinity Large Thinking score almost the same on the Noometry Index (39.4 vs 38.6), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Instruct or Trinity Large Thinking better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 34.1 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Trinity Large Thinking share?

16 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper