Arcee AI, open weights
Trinity Large Thinking
Trinity Large Thinking by Arcee AI ranks 185th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is knowledge, where it ranks 113th. API pricing starts at $0.25 per million input tokens and $0.80 per million output tokens, with a 262K-token context window.
Last verified
Specifications
- Noometry rank
- #185 of 354
- Index score
- 38.6
- Evidence
- Confirmed 24 results
- Provider
Arcee AI
- Released
- April 1, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 262K
- Max output
- 80K
- Input price
- $0.25 / M
- Output price
- $0.80 / M
- Blended price
- $0.39 / M
- Output speed
- Not measured
- Value
- #61 of 219
- Knowledge cutoff
- Unknown
- Input
- text
- Hugging Face
- arcee-ai/Trinity-Large-Thinking
Category scores
Each category score combines every public result we have in that category.
- Coding 34.1
- Reasoning 16.9
- Math 37.6
- Knowledge 40.9
- Multilingual 46.2
- Instruction Following 70.5
- Long Context 41.3
- Writing & Preference 53.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.1 | #244 | 3 |
| Reasoning | 16.9 | #298 | 5 |
| Math | 37.6 | #149 | 1 |
| Knowledge | 40.9 | #113 | 2 |
| Multilingual | 46.2 | #160 | 1 |
| Instruction Following | 70.5 | #162 | 1 |
| Long Context | 41.3 | #144 | 1 |
| Writing & Preference | 53.8 | #158 | 3 |
Strengths and weaknesses
Categories where Trinity Large Thinking places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Knowledge | 40.9 | +3.6 | #113 of 314, top 36% |
| Math | 37.6 | +1.1 | #149 of 327, top 46% |
| Long Context | 41.3 | +0.3 | #144 of 296, top 49% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 16.9 | −6.7 | #298 of 350, top 86% |
| Coding | 34.1 | −4.6 | #244 of 340, top 72% |
| Multilingual | 46.2 | −1.2 | #160 of 297, top 54% |
Closest competitors
The models ranked just above and below Trinity Large Thinking. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen2.5 Plus 1127 | #181 | 38.8 | — | — | Compare |
| Qwen3.6 Flash | #182 | 38.8 | $0.42 | — | Compare |
| Olmo 3 32b Think | #183 | 38.7 | — | — | Compare |
| Hunyuan Large 2025 02 10 | #184 | 38.6 | — | — | Compare |
| GPT-5.1-Codex | #186 | 38.6 | $3.44 | — | Compare |
| Sonar | #187 | 38.5 | $1 | — | Compare |
| MiniMax-M2.5 | #188 | 38.3 | $0.52 | 49 | Compare |
| Nova Premier 1.0 | #189 | 38.3 | $5 | 9 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena WebDev | 1238 | #105 of 113, top 93% | thinking | LMArena | 2026-10-08 |
| SciCode | 36.1% | #92 of 121, top 77% | Epoch AI | ||
| LMArena Coding | 1381 | #147 of 294, top 50% | LMArena | 2026-10-08 | |
| LMArena Coding | 1357 | thinking | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| NYT Connections (extended) | 16.5% | #81 of 91, top 90% | Lech Mazur benchmarks | ||
| CritPt | 0.9% | #86 of 134, top 65% | Epoch AI | ||
| Thematic Generalization | 41.6% | #20 of 23, top 87% | Lech Mazur benchmarks | ||
| LMArena Hard Prompts | 1350 | #160 of 297, top 54% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1342 | thinking | LMArena | 2026-10-08 | |
| Surface Evolver Bench | 15.6% | #25 of 25, top 100% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1336 | LMArena | 2026-10-08 | ||
| LMArena Math | 1366 | #152 of 285, top 54% | thinking | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Vectara Hallucination Rate (lower is better) | 6.9% | #27 of 96, top 29% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1352 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1360 | #148 of 273, top 55% | thinking | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1318 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1325 | #160 of 297, top 54% | thinking | LMArena | 2026-10-08 |
| LMArena Chinese | 1352 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1373 | #155 of 285, top 55% | thinking | LMArena | 2026-10-08 |
| LMArena French | 1374 | #133 of 223, top 60% | LMArena | 2026-10-08 | |
| LMArena French | 1365 | thinking | LMArena | 2026-10-08 | |
| LMArena German | 1313 | LMArena | 2026-10-08 | ||
| LMArena German | 1356 | #124 of 231, top 54% | thinking | LMArena | 2026-10-08 |
| LMArena Japanese | 1290 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1311 | #115 of 211, top 55% | thinking | LMArena | 2026-10-08 |
| LMArena Korean | 1258 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1306 | #128 of 213, top 61% | thinking | LMArena | 2026-10-08 |
| LMArena Russian | 1325 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1337 | #153 of 283, top 55% | thinking | LMArena | 2026-10-08 |
| LMArena Spanish | 1357 | #139 of 226, top 62% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1328 | thinking | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1334 | #155 of 298, top 53% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1325 | thinking | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1355 | #150 of 291, top 52% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1318 | thinking | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1339 | LMArena | 2026-10-08 | ||
| LMArena Text | 1340 | #164 of 297, top 56% | thinking | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1320 | #152 of 295, top 52% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1307 | thinking | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1342 | #158 of 295, top 54% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1320 | thinking | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| openrouter | $0.25 | $0.80 | $0.06 | 2026-10-10 |
Compare Trinity Large Thinking
- Trinity Large Thinking vs Hunyuan Large 2025 02 10
- Trinity Large Thinking vs GPT-5.1-Codex
- Trinity Large Thinking vs Olmo 3 32b Think
- Trinity Large Thinking vs Sonar
- Trinity Large Thinking vs Qwen3.6 Flash
- Trinity Large Thinking vs MiniMax-M2.5
- Trinity Large Thinking vs GPT-6 Astra
- Trinity Large Thinking vs Claude Fable 5.1
- Trinity Large Thinking vs Gemini 3.8 Flash
- Trinity Large Thinking vs Kimi K3
- Trinity Large Thinking vs Grok 4.6
- Trinity Large Thinking vs Qwen3.8 Max
- Trinity Large Thinking vs GLM-5.3
- Trinity Large Thinking vs Muse Spark 1.3
Frequently asked questions
How good is Trinity Large Thinking?
Trinity Large Thinking by Arcee AI ranks 185th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is knowledge, where it ranks 113th. API pricing starts at $0.25 per million input tokens and $0.80 per million output tokens, with a 262K-token context window.
How much does Trinity Large Thinking cost?
Trinity Large Thinking costs $0.25 per million input tokens and $0.80 per million output tokens on openrouter, with cached input at $0.06.
What is Trinity Large Thinking's context window?
Trinity Large Thinking accepts up to 262K tokens of input and can write up to 80K tokens in one response.
Is Trinity Large Thinking open source?
Yes. Trinity Large Thinking's weights are downloadable from Hugging Face (arcee-ai/Trinity-Large-Thinking); check the license for commercial terms.
What are Trinity Large Thinking's strengths and weaknesses?
Relative to other ranked models, Trinity Large Thinking places best in knowledge, math, long context and lowest in reasoning, coding, multilingual.
What is Trinity Large Thinking best at?
Its best category is knowledge, where it ranks 113th on Noometry.