Model comparison

Gemma 2B vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 29.6 on the Noometry Index.

Last verified . 11 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Gemma 2B scores higher in 1 category and Trinity Large Thinking in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Trinity Large Thinking leads 53.8 to 24.0.

Side by side

Gemma 2B and Trinity Large Thinking specifications
Gemma 2BTrinity Large Thinking
ProviderGoogleArcee AI
Noometry Index29.638.6
Released2024-02-212026-04-01
WeightsOpenOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked2324

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Trinity Large Thinking leads

Gemma 2B: 29.4 (#305), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Coding10101381
LMArena WebDev—1238
SciCode—36.1%
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Gemma 2B leads

Gemma 2B: 18.8 (#275), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Hard Prompts9891350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Trinity Large Thinking leads

Gemma 2B: 30.0 (#239), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Math10091366
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGemma 2BTrinity Large Thinking
Vectara Hallucination Rate—6.9%
LMArena Expert—1360
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multilingual Trinity Large Thinking leads

Gemma 2B: 23.0 (#294), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Non-English9581325
LMArena Chinese9861373
LMArena Russian9371337
LMArena French—1374
LMArena German—1356
LMArena Japanese—1311
LMArena Korean—1306
LMArena Spanish—1357

Instruction Following Trinity Large Thinking leads

Gemma 2B: 48.5 (#302), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Instruction Following9701334

Long Context Trinity Large Thinking leads

Gemma 2B: 29.9 (#291), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Longer Query9811355

Writing & Preference Trinity Large Thinking leads

Gemma 2B: 24.0 (#308), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGemma 2BTrinity Large Thinking
LMArena Text10021340
LMArena Creative Writing9871320
LMArena Multi-Turn9451342

Frequently asked questions

Is Gemma 2B better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 29.6 on the Noometry Index.

Is Gemma 2B or Trinity Large Thinking better for coding?

Trinity Large Thinking scores higher on coding benchmarks: 34.1 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and Trinity Large Thinking share?

11 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper