Model comparison

Mixtral 8x7B vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 27.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mixtral 8x7B scores higher in 1 category and Trinity Large Thinking in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Trinity Large Thinking leads 40.9 to 11.0.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $0.70 / $0.70 for Mixtral 8x7B.
  • Trinity Large Thinking accepts more context: 262K tokens versus 32K.

Side by side

Mixtral 8x7B and Trinity Large Thinking specifications
Mixtral 8x7BTrinity Large Thinking
ProviderMistral AIArcee AI
Noometry Index27.138.6
Released2023-12-112026-04-01
WeightsOpenOpen
Context window32K262K
Max output32K80K
Input $ / M tokens$0.70$0.25
Output $ / M tokens$0.70$0.80
Results tracked3824

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Trinity Large Thinking leads

Mixtral 8x7B: 32.8 (#269), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Coding11261381
LMArena WebDev—1238
SciCode—36.1%
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Mixtral 8x7B leads

Mixtral 8x7B: 18.2 (#285), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Hard Prompts11151350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
DTBench49.6%—
Surface Evolver Bench—15.6%
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Trinity Large Thinking leads

Mixtral 8x7B: 18.8 (#289), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Math11471366
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Trinity Large Thinking leads

Mixtral 8x7B: 11.0 (#301), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Expert10881360
GPQA Diamond30.6%—
MMLU-Pro33.5%—
Vectara Hallucination Rate—6.9%
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Trinity Large Thinking leads

Mixtral 8x7B: 29.6 (#266), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Non-English10771325
LMArena Chinese10551373
LMArena French11661374
LMArena German11141356
LMArena Japanese9311311
LMArena Korean9681306
LMArena Russian10901337
LMArena Spanish11111357

Instruction Following Trinity Large Thinking leads

Mixtral 8x7B: 51.0 (#297), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Instruction Following11091334
IFEval57.5%—

Long Context Trinity Large Thinking leads

Mixtral 8x7B: 33.4 (#260), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Longer Query11031355

Writing & Preference Trinity Large Thinking leads

Mixtral 8x7B: 34.2 (#270), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BTrinity Large Thinking
LMArena Text11321340
LMArena Creative Writing11091320
LMArena Multi-Turn11151342
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 27.1 on the Noometry Index.

Which is cheaper, Mixtral 8x7B or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Mixtral 8x7B lists at $0.70 and $0.70.

Is Mixtral 8x7B or Trinity Large Thinking better for coding?

Trinity Large Thinking scores higher on coding benchmarks: 34.1 versus 32.8 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 32K.

How many benchmarks do Mixtral 8x7B and Trinity Large Thinking share?

17 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper