Model comparison

Mistral vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 29.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral scores higher in 1 category and Trinity Large Thinking in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Trinity Large Thinking leads 40.9 to 16.6.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Mistral and Trinity Large Thinking specifications
MistralTrinity Large Thinking
ProviderMistral AIArcee AI
Noometry Index29.938.6
Released—2026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral: 33.8 (#250), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Coding11621381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Mistral leads

Mistral: 22.2 (#200), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Hard Prompts11491350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Mistral: 22.3 (#278), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Math11801366
Omni-MATH7.2%—

Knowledge Trinity Large Thinking leads

Mistral: 16.6 (#288), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Expert11251360
MMLU-Pro27.7%—
Vectara Hallucination Rate—6.9%
GPQA (HELM)30.3%—

Multilingual Trinity Large Thinking leads

Mistral: 32.8 (#254), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Non-English11291325
LMArena Chinese11091373
LMArena French11801374
LMArena German11551356
LMArena Japanese10131311
LMArena Korean10321306
LMArena Russian11681337
LMArena Spanish11431357

Instruction Following Trinity Large Thinking leads

Mistral: 52.6 (#288), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Instruction Following11521334
IFEval56.8%—

Long Context Trinity Large Thinking leads

Mistral: 35.0 (#245), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Longer Query11531355

Writing & Preference Trinity Large Thinking leads

Mistral: 37.0 (#260), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMistralTrinity Large Thinking
LMArena Text11651340
LMArena Creative Writing11581320
LMArena Multi-Turn11471342
WildBench66%—

Frequently asked questions

Is Mistral better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 29.9 on the Noometry Index.

Is Mistral or Trinity Large Thinking better for coding?

They score almost the same on coding (33.8 vs 34.1); test both on your own repository before choosing.

How many benchmarks do Mistral and Trinity Large Thinking share?

17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper