Model comparison

Mistral Medium 3.5 vs Trinity Large Thinking

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 38.6 on the Noometry Index. Trinity Large Thinking costs 7.7× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Last verified . 18 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 7 categories and Trinity Large Thinking in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Medium 3.5 leads 51.9 to 46.2.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.

Side by side

Mistral Medium 3.5 and Trinity Large Thinking specifications
Mistral Medium 3.5Trinity Large Thinking
ProviderMistral AIArcee AI
Noometry Index40.238.6
Released—2026-04-01
WeightsOpenOpen
Context window262K262K
Max output210K80K
Input $ / M tokens$1.50$0.25
Output $ / M tokens$7.50$0.80
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Mistral Medium 3.5: 36.0 (#213), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena WebDev12641238
LMArena Coding14611381
SciCode—36.1%

Reasoning Too close to call

Mistral Medium 3.5: 17.3 (#295), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
NYT Connections (extended)12.9%16.5%
LMArena Hard Prompts14361350
Kagi LLM Benchmark41.4%—
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%
Epoch Capabilities Index141.35—

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Math14311366

Knowledge Too close to call

Mistral Medium 3.5: 40.0 (#126), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Expert14321360
Vectara Hallucination Rate—6.9%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Non-English14041325
LMArena Chinese14421373
LMArena French14481374
LMArena German14511356
LMArena Korean13851306
LMArena Russian13951337
LMArena Spanish14091357
LMArena Japanese—1311

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Instruction Following14151334

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Longer Query14151355

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Trinity Large Thinking
LMArena Text14211340
LMArena Creative Writing13741320
LMArena Multi-Turn14231342
EQ-Bench 4993—

Frequently asked questions

Is Mistral Medium 3.5 better than Trinity Large Thinking?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 38.6 on the Noometry Index. Trinity Large Thinking costs 7.7× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.5 or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Trinity Large Thinking better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Mistral Medium 3.5 and Trinity Large Thinking share?

18 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper