Model comparison

Mistral Medium vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 36.3 on the Noometry Index.

Last verified . 21 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Mistral Medium scores higher in 6 categories and Trinity Large Thinking in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Trinity Large Thinking leads 40.9 to 25.0.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 22.7% for Mistral Medium and 6.9% for Trinity Large Thinking.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.

Side by side

Mistral Medium and Trinity Large Thinking specifications
Mistral MediumTrinity Large Thinking
ProviderMistral AIArcee AI
Noometry Index36.338.6
Released2023-12-112026-04-01
WeightsOpenOpen
Context window262K262K
Max output262K80K
Input $ / M tokens$1.50$0.25
Output $ / M tokens$7.50$0.80
Results tracked3624

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Medium: 34.2 (#243), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMistral MediumTrinity Large Thinking
SciCode40.2%36.1%
LMArena Coding14341381
FrontierCode8%—
LMArena WebDev—1238
WeirdML43.7%—
ALE-Bench763.98—

Agentic & Tool Use Not comparable

Mistral Medium: 28.3 (#90), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkMistral MediumTrinity Large Thinking
Berkeley Function Calling Leaderboard37.7%—

Reasoning Mistral Medium leads

Mistral Medium: 24.0 (#167), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMistral MediumTrinity Large Thinking
CritPt0%0.9%
LMArena Hard Prompts14261350
Surface Evolver Bench26.9%15.6%
Kagi LLM Benchmark50%—
NYT Connections (extended)—16.5%
Thematic Generalization—41.6%
DTBench75.5%—
LMCA26.1%—

Math Trinity Large Thinking leads

Mistral Medium: 28.1 (#245), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMistral MediumTrinity Large Thinking
LMArena Math14081366
OTIS Mock AIME 2024-202532.2%—
ProofBench9%—
MATH Level 581.6%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Trinity Large Thinking leads

Mistral Medium: 25.0 (#265), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMistral MediumTrinity Large Thinking
Vectara Hallucination Rate22.7%6.9%
LMArena Expert14081360
GPQA Diamond59.5%—
Humanity's Last Exam4.5%—

Multimodal Not comparable

Mistral Medium: 35.3 (#88), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkMistral MediumTrinity Large Thinking
LMArena Vision1172—

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMistral MediumTrinity Large Thinking
LMArena Non-English14081325
LMArena Chinese14471373
LMArena French14591374
LMArena German14321356
LMArena Japanese13781311
LMArena Korean13801306
LMArena Russian14111337
LMArena Spanish14331357

Instruction Following Mistral Medium leads

Mistral Medium: 73.7 (#116), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMistral MediumTrinity Large Thinking
LMArena Instruction Following13981334

Long Context Mistral Medium leads

Mistral Medium: 42.9 (#114), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMistral MediumTrinity Large Thinking
LMArena Longer Query14061355

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMistral MediumTrinity Large Thinking
LMArena Text14241340
LMArena Creative Writing13911320
LMArena Multi-Turn14181342
Short-Story Creative Writing77.3%—

Frequently asked questions

Is Mistral Medium better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 36.3 on the Noometry Index.

Which is cheaper, Mistral Medium or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mistral Medium or Trinity Large Thinking better for coding?

They score almost the same on coding (34.2 vs 34.1); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Mistral Medium and Trinity Large Thinking share?

21 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper