Model comparison

DeepSeek-V3 vs Trinity Large Thinking

DeepSeek-V3 and Trinity Large Thinking score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 20 benchmarks with published results for both. DeepSeek-V3 scores higher in 5 categories and Trinity Large Thinking in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where DeepSeek-V3 leads 42.3 to 34.1.
  • Both cost about the same: $0.24 input and $0.90 output per million tokens.
  • Trinity Large Thinking accepts more context: 262K tokens versus 164K.

Side by side

DeepSeek-V3 and Trinity Large Thinking specifications
DeepSeek-V3Trinity Large Thinking
ProviderDeepSeekArcee AI
Noometry Index39.538.6
Released2024-12-262026-04-01
WeightsOpenOpen
Context window164K262K
Max output164K80K
Input $ / M tokens$0.24$0.25
Output $ / M tokens$0.90$0.80
Results tracked6024

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
SciCode35.8%36.1%
LMArena Coding13681381
Aider Polyglot55.1%—
LMArena WebDev—1238
WeirdML36.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
METR Time Horizons49.6%—

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
CritPt0%0.9%
LMArena Hard Prompts13651350
SimpleBench27.2%—
Kagi LLM Benchmark52.3%—
NYT Connections (extended)—16.5%
Thematic Generalization—41.6%
LiveBench Reasoning65.8%—
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
Surface Evolver Bench—15.6%
BIG-Bench Hard87.5%—
Epoch Capabilities Index135.94—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Trinity Large Thinking leads

DeepSeek-V3: 32.1 (#219), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
LMArena Math13731366
OTIS Mock AIME 2024-202537.8%—
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Trinity Large Thinking leads

DeepSeek-V3: 37.5 (#155), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
Vectara Hallucination Rate6.1%6.9%
LMArena Expert13511360
GPQA Diamond67.6%—
MMLU-Pro72.3%—
Confabulations26.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
LMArena Non-English13581325
LMArena Chinese13911373
LMArena French13851374
LMArena German13741356
LMArena Japanese13331311
LMArena Korean13191306
LMArena Russian13731337
LMArena Spanish13581357

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
LMArena Instruction Following13451334
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context Trinity Large Thinking leads

DeepSeek-V3: 34.0 (#253), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
LMArena Longer Query13521355
Fiction.LiveBench50%—

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Trinity Large Thinking
LMArena Text13751340
LMArena Creative Writing13641320
LMArena Multi-Turn13891342
Short-Story Creative Writing77%—
EQ-Bench Creative Writing1472—
WildBench83%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Trinity Large Thinking?

DeepSeek-V3 and Trinity Large Thinking score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3 or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; DeepSeek-V3 lists at $0.24 and $0.90.

Is DeepSeek-V3 or Trinity Large Thinking better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-V3 and Trinity Large Thinking share?

20 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper