Model comparison
Mistral Medium 3.1 vs Trinity Large Thinking
Trinity Large Thinking is the stronger model overall, scoring 38.6 to 31.9 on the Noometry Index.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. Mistral Medium 3.1 scores higher in 1 category and Trinity Large Thinking in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Trinity Large Thinking leads 16.9 to 10.6.
- The biggest single-benchmark swing is Thematic Generalization: 20.3% for Mistral Medium 3.1 and 41.6% for Trinity Large Thinking.
- Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $0.40 / $2 for Mistral Medium 3.1.
- Trinity Large Thinking accepts more context: 262K tokens versus 131K.
- Trinity Large Thinking has downloadable open weights; the other is API-only.
Side by side
| Mistral Medium 3.1 | Trinity Large Thinking | |
|---|---|---|
| Provider | Mistral AI | Arcee AI |
| Noometry Index | 31.9 | 38.6 |
| Released | — | 2026-04-01 |
| Weights | Proprietary | Open |
| Context window | 131K | 262K |
| Max output | 105K | 80K |
| Input $ / M tokens | $0.40 | $0.25 |
| Output $ / M tokens | $2 | $0.80 |
| Results tracked | 3 | 24 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Mistral Medium 3.1: —, Trinity Large Thinking: 34.1 (#244)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| LMArena WebDev | — | 1238 |
| SciCode | — | 36.1% |
| LMArena Coding | — | 1381 |
Reasoning Trinity Large Thinking leads
Mistral Medium 3.1: 10.6 (#341), Trinity Large Thinking: 16.9 (#298)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| NYT Connections (extended) | 6.5% | 16.5% |
| Thematic Generalization | 20.3% | 41.6% |
| CritPt | — | 0.9% |
| LMArena Hard Prompts | — | 1350 |
| Surface Evolver Bench | — | 15.6% |
Math Not comparable
Mistral Medium 3.1: —, Trinity Large Thinking: 37.6 (#149)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| LMArena Math | — | 1366 |
Knowledge Not comparable
Mistral Medium 3.1: —, Trinity Large Thinking: 40.9 (#113)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| Vectara Hallucination Rate | — | 6.9% |
| LMArena Expert | — | 1360 |
Multilingual Not comparable
Mistral Medium 3.1: —, Trinity Large Thinking: 46.2 (#160)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| LMArena Non-English | — | 1325 |
| LMArena Chinese | — | 1373 |
| LMArena French | — | 1374 |
| LMArena German | — | 1356 |
| LMArena Japanese | — | 1311 |
| LMArena Korean | — | 1306 |
| LMArena Russian | — | 1337 |
| LMArena Spanish | — | 1357 |
Instruction Following Not comparable
Mistral Medium 3.1: —, Trinity Large Thinking: 70.5 (#162)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| LMArena Instruction Following | — | 1334 |
Long Context Not comparable
Mistral Medium 3.1: —, Trinity Large Thinking: 41.3 (#144)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| LMArena Longer Query | — | 1355 |
Writing & Preference Mistral Medium 3.1 leads
Mistral Medium 3.1: 55.5 (#145), Trinity Large Thinking: 53.8 (#158)
| Benchmark | Mistral Medium 3.1 | Trinity Large Thinking |
|---|---|---|
| LMArena Text | — | 1340 |
| LMArena Creative Writing | — | 1320 |
| EQ-Bench Creative Writing | 1476 | — |
| LMArena Multi-Turn | — | 1342 |
Frequently asked questions
Is Mistral Medium 3.1 better than Trinity Large Thinking?
Trinity Large Thinking is the stronger model overall, scoring 38.6 to 31.9 on the Noometry Index.
Which is cheaper, Mistral Medium 3.1 or Trinity Large Thinking?
Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Mistral Medium 3.1 lists at $0.40 and $2.
Which has the bigger context window?
Trinity Large Thinking does, with 262K tokens against 131K.
How many benchmarks do Mistral Medium 3.1 and Trinity Large Thinking share?
2 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Trinity Large Thinking has 24.