Model comparison

Gemini 1.5 Flash 8B vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 29.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 2 categories and Trinity Large Thinking in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Trinity Large Thinking leads 40.9 to 16.0.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash 8B and Trinity Large Thinking specifications
Gemini 1.5 Flash 8BTrinity Large Thinking
ProviderGoogleArcee AI
Noometry Index29.938.6
Released2024-10-032026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 35.5 (#225), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Coding12181381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 20.0 (#244), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Hard Prompts12091350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
DTBench50%—
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Gemini 1.5 Flash 8B: 14.2 (#302), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Math12071366
OTIS Mock AIME 2024-20254.6%—

Knowledge Trinity Large Thinking leads

Gemini 1.5 Flash 8B: 16.0 (#289), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Expert11851360
GPQA Diamond33%—
Vectara Hallucination Rate—6.9%

Multimodal Not comparable

Gemini 1.5 Flash 8B: 28.2 (#115), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Vision1044—

Multilingual Trinity Large Thinking leads

Gemini 1.5 Flash 8B: 38.5 (#229), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Non-English12151325
LMArena Chinese12311373
LMArena French12341374
LMArena German12061356
LMArena Japanese11501311
LMArena Korean11401306
LMArena Russian12361337
LMArena Spanish12121357

Instruction Following Trinity Large Thinking leads

Gemini 1.5 Flash 8B: 62.8 (#236), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Instruction Following11991334

Long Context Trinity Large Thinking leads

Gemini 1.5 Flash 8B: 37.0 (#225), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Longer Query12191355

Writing & Preference Trinity Large Thinking leads

Gemini 1.5 Flash 8B: 42.8 (#232), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BTrinity Large Thinking
LMArena Text12261340
LMArena Creative Writing12181320
LMArena Multi-Turn11851342

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or Trinity Large Thinking better for coding?

Gemini 1.5 Flash 8B scores higher on coding benchmarks: 35.5 versus 34.1 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and Trinity Large Thinking share?

17 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper