Model comparison

Grok 4.1 Fast vs Trinity Large Thinking

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 38.6 on the Noometry Index.

Last verified . 20 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Grok 4.1 Fast scores higher in 5 categories and Trinity Large Thinking in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 16.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 87.4% for Grok 4.1 Fast and 16.5% for Trinity Large Thinking.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $0.25 / $0.80 for Trinity Large Thinking.
  • Trinity Large Thinking accepts more context: 262K tokens versus 128K.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 Fast and Trinity Large Thinking specifications
Grok 4.1 FastTrinity Large Thinking
ProviderxAIArcee AI
Noometry Index41.438.6
Released2025-06-272026-04-01
WeightsProprietaryOpen
Context window128K262K
Max output30K80K
Input $ / M tokens$0.20$0.25
Output $ / M tokens$0.50$0.80
Results tracked3224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1 Fast: 34.1 (#245), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena WebDev12421238
LMArena Coding14111381
SciCode—36.1%
ALE-Bench394.93—

Agentic & Tool Use Not comparable

Grok 4.1 Fast: 36.3 (#39), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
NYT Connections (extended)87.4%16.5%
LMArena Hard Prompts14071350
SimpleBench56%—
CritPt—0.9%
Thematic Generalization—41.6%
DTBench87.7%—
Surface Evolver Bench—15.6%
ForecastBench61—

Math Trinity Large Thinking leads

Grok 4.1 Fast: 31.9 (#221), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena Math14081366
MathArena Final-Answer Competitions60.9%—
ProofBench4%—

Knowledge Trinity Large Thinking leads

Grok 4.1 Fast: 33.1 (#207), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
Vectara Hallucination Rate17.8%6.9%
LMArena Expert13991360

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena Vision1201—

Multilingual Grok 4.1 Fast leads

Grok 4.1 Fast: 51.0 (#114), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena Non-English13911325
LMArena Chinese14411373
LMArena French14151374
LMArena German14041356
LMArena Japanese13491311
LMArena Korean13611306
LMArena Russian13871337
LMArena Spanish14131357

Instruction Following Grok 4.1 Fast leads

Grok 4.1 Fast: 72.7 (#133), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena Instruction Following13761334

Long Context Grok 4.1 Fast leads

Grok 4.1 Fast: 42.4 (#126), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena Longer Query13901355

Writing & Preference Grok 4.1 Fast leads

Grok 4.1 Fast: 57.2 (#131), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastTrinity Large Thinking
LMArena Text14081340
LMArena Creative Writing13941320
LMArena Multi-Turn13891342
EQ-Bench Creative Writing1327—

Frequently asked questions

Is Grok 4.1 Fast better than Trinity Large Thinking?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 38.6 on the Noometry Index.

Which is cheaper, Grok 4.1 Fast or Trinity Large Thinking?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; Trinity Large Thinking lists at $0.25 and $0.80.

Is Grok 4.1 Fast or Trinity Large Thinking better for coding?

They score almost the same on coding (34.1 vs 34.1); test both on your own repository before choosing.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 128K.

How many benchmarks do Grok 4.1 Fast and Trinity Large Thinking share?

20 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper