Model comparison

Grok-2 (Dec 2024) vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 33.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok-2 (Dec 2024) scores higher in 0 categories and Trinity Large Thinking in 8 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Trinity Large Thinking leads 37.6 to 20.8.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Grok-2 (Dec 2024) and Trinity Large Thinking specifications
Grok-2 (Dec 2024)Trinity Large Thinking
ProviderxAIArcee AI
Noometry Index33.738.6
Released2024-08-132026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked3424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok-2 (Dec 2024): 33.3 (#258), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Coding12871381
LMArena WebDev—1238
SciCode—36.1%
WeirdML22.2%—
LiveBench Coding46.4%—

Reasoning Too close to call

Grok-2 (Dec 2024): 16.9 (#299), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Hard Prompts12721350
SimpleBench22.7%—
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
LiveBench Reasoning54.8%—
DTBench65.2%—
LiveBench Data Analysis54.5%—
Surface Evolver Bench—15.6%
Epoch Capabilities Index130.48—
LiveBench54.3%—

Math Trinity Large Thinking leads

Grok-2 (Dec 2024): 20.8 (#284), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Math12831366
OTIS Mock AIME 2024-202511.5%—
LiveBench Math54.9%—
MATH Level 563.5%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Trinity Large Thinking leads

Grok-2 (Dec 2024): 29.8 (#233), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Expert12541360
GPQA Diamond53.8%—
Confabulations20.1%—
Vectara Hallucination Rate—6.9%

Multilingual Trinity Large Thinking leads

Grok-2 (Dec 2024): 43.1 (#188), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Non-English12821325
LMArena Chinese12891373
LMArena French13181374
LMArena German12871356
LMArena Japanese12441311
LMArena Korean12371306
LMArena Russian12861337
LMArena Spanish12811357

Instruction Following Trinity Large Thinking leads

Grok-2 (Dec 2024): 66.9 (#202), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Instruction Following12701334
LiveBench Instruction Following69.6%—

Long Context Trinity Large Thinking leads

Grok-2 (Dec 2024): 38.8 (#190), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Longer Query12761355

Writing & Preference Trinity Large Thinking leads

Grok-2 (Dec 2024): 48.6 (#198), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGrok-2 (Dec 2024)Trinity Large Thinking
LMArena Text13051340
LMArena Creative Writing12841320
LMArena Multi-Turn12901342
Short-Story Creative Writing63.6%—
LiveBench Language45.6%—

Frequently asked questions

Is Grok-2 (Dec 2024) better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 33.7 on the Noometry Index.

Is Grok-2 (Dec 2024) or Trinity Large Thinking better for coding?

They score almost the same on coding (33.3 vs 34.1); test both on your own repository before choosing.

How many benchmarks do Grok-2 (Dec 2024) and Trinity Large Thinking share?

17 benchmarks have published results for both models. Grok-2 (Dec 2024) has 34 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper