Model comparison

Grok 4.1 vs Trinity Large Thinking

Grok 4.1 is the stronger model overall, scoring 41.5 to 38.6 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.1 scores higher in 6 categories and Trinity Large Thinking in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 leads 29.5 to 16.9.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Trinity Large Thinking specifications
Grok 4.1Trinity Large Thinking
ProviderxAIArcee AI
Noometry Index41.538.6
Released2025-11-172026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked1924

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1: 33.7 (#253), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena WebDev12141238
LMArena Coding14451381
SciCode—36.1%

Agentic & Tool Use Not comparable

Grok 4.1: 34.1 (#49), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
Cybench39%—

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Hard Prompts14351350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Grok 4.1 leads

Grok 4.1: 38.9 (#120), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Math14221366

Knowledge Trinity Large Thinking leads

Grok 4.1: 39.5 (#133), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Expert14171360
Vectara Hallucination Rate—6.9%

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Non-English14251325
LMArena Chinese14651373
LMArena French14481374
LMArena German14461356
LMArena Japanese13971311
LMArena Korean14071306
LMArena Russian14341337
LMArena Spanish14381357

Instruction Following Grok 4.1 leads

Grok 4.1: 73.8 (#111), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Instruction Following14001334

Long Context Grok 4.1 leads

Grok 4.1: 43.2 (#100), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Longer Query14161355

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGrok 4.1Trinity Large Thinking
LMArena Text14371340
LMArena Creative Writing14111320
LMArena Multi-Turn14371342

Frequently asked questions

Is Grok 4.1 better than Trinity Large Thinking?

Grok 4.1 is the stronger model overall, scoring 41.5 to 38.6 on the Noometry Index.

Is Grok 4.1 or Trinity Large Thinking better for coding?

They score almost the same on coding (33.7 vs 34.1); test both on your own repository before choosing.

How many benchmarks do Grok 4.1 and Trinity Large Thinking share?

18 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper