Model comparison

Magistral Medium vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 35.2 on the Noometry Index.

Last verified . 19 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Magistral Medium scores higher in 1 category and Trinity Large Thinking in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Trinity Large Thinking leads 16.9 to 8.6.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $2 / $5 for Magistral Medium.

Side by side

Magistral Medium and Trinity Large Thinking specifications
Magistral MediumTrinity Large Thinking
ProviderMistral AIArcee AI
Noometry Index35.238.6
Released2025-03-172026-04-01
WeightsOpenOpen
Context window262K262K
Max output16K80K
Input $ / M tokens$2$0.25
Output $ / M tokens$5$0.80
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Magistral Medium: 39.1 (#161), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
SciCode39.2%36.1%
LMArena Coding13191381
LMArena WebDev—1238

Reasoning Trinity Large Thinking leads

Magistral Medium: 8.6 (#348), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
CritPt0.3%0.9%
LMArena Hard Prompts12671350
ARC-AGI-20%—
Kagi LLM Benchmark16.2%—
NYT Connections (extended)—16.5%
ARC-AGI-16.1%—
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Magistral Medium: 35.1 (#189), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
LMArena Math12501366

Knowledge Trinity Large Thinking leads

Magistral Medium: 33.5 (#202), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
LMArena Expert12231360
Vectara Hallucination Rate—6.9%

Multilingual Trinity Large Thinking leads

Magistral Medium: 39.6 (#224), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
LMArena Non-English12321325
LMArena Chinese12271373
LMArena French12671374
LMArena German12481356
LMArena Japanese11751311
LMArena Korean11251306
LMArena Russian12241337
LMArena Spanish12711357

Instruction Following Trinity Large Thinking leads

Magistral Medium: 66.0 (#211), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
LMArena Instruction Following12541334

Long Context Trinity Large Thinking leads

Magistral Medium: 39.3 (#183), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
LMArena Longer Query12951355

Writing & Preference Trinity Large Thinking leads

Magistral Medium: 46.3 (#219), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMagistral MediumTrinity Large Thinking
LMArena Text12551340
LMArena Creative Writing12451320
LMArena Multi-Turn12751342

Frequently asked questions

Is Magistral Medium better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 35.2 on the Noometry Index.

Which is cheaper, Magistral Medium or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Magistral Medium lists at $2 and $5.

Is Magistral Medium or Trinity Large Thinking better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Magistral Medium and Trinity Large Thinking share?

19 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper