Model comparison

Qwen2.5 72B Instruct vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Qwen2.5 72B Instruct scores higher in 1 category and Trinity Large Thinking in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Trinity Large Thinking leads 37.6 to 19.3.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $1.40 / $5.60 for Qwen2.5 72B Instruct.
  • Trinity Large Thinking accepts more context: 262K tokens versus 131K.

Side by side

Qwen2.5 72B Instruct and Trinity Large Thinking specifications
Qwen2.5 72B InstructTrinity Large Thinking
ProviderAlibaba (Qwen)Arcee AI
Noometry Index31.938.6
Released2024-092026-04-01
WeightsOpenOpen
Context window131K262K
Max output8K80K
Input $ / M tokens$1.40$0.25
Output $ / M tokens$5.60$0.80
Results tracked4324

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Qwen2.5 72B Instruct: 33.2 (#260), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Coding12921381
LMArena WebDev—1238
SciCode—36.1%
WeirdML16%—
BigCodeBench Instruct45.8%—
BigCodeBench Complete55.9%—

Agentic & Tool Use Not comparable

Qwen2.5 72B Instruct: 22.1 (#133), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
TheAgentCompany5.7%—
BALROG16.2%—
METR Time Horizons35.8%—

Reasoning Qwen2.5 72B Instruct leads

Qwen2.5 72B Instruct: 22.3 (#199), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Hard Prompts12711350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
DTBench62.9%—
LMCA13.4%—
Surface Evolver Bench—15.6%
BIG-Bench Hard79.8%—
Epoch Capabilities Index129—
ForecastBench57.5—
HellaSwag84.8%—
PIQA82.6%—
WinoGrande82.3%—

Math Trinity Large Thinking leads

Qwen2.5 72B Instruct: 19.3 (#287), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Math12831366
OTIS Mock AIME 2024-20258.1%—
Omni-MATH33%—
MATH Level 563.2%—

Knowledge Trinity Large Thinking leads

Qwen2.5 72B Instruct: 27.0 (#253), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Expert12451360
GPQA Diamond49.1%—
MMLU-Pro63.1%—
Confabulations19.1%—
Vectara Hallucination Rate—6.9%
GPQA (HELM)42.6%—
ARC (AI2) Challenge94.5%—
MMLU85.3%—
TriviaQA71.9%—

Multilingual Trinity Large Thinking leads

Qwen2.5 72B Instruct: 41.0 (#213), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Non-English12521325
LMArena Chinese12721373
LMArena French12801374
LMArena German12341356
LMArena Japanese11801311
LMArena Korean11881306
LMArena Russian12641337
LMArena Spanish12561357

Instruction Following Trinity Large Thinking leads

Qwen2.5 72B Instruct: 65.5 (#221), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Instruction Following12541334
IFEval80.6%—

Long Context Trinity Large Thinking leads

Qwen2.5 72B Instruct: 38.9 (#188), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Longer Query12821355

Writing & Preference Trinity Large Thinking leads

Qwen2.5 72B Instruct: 46.7 (#215), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkQwen2.5 72B InstructTrinity Large Thinking
LMArena Text12691340
LMArena Creative Writing12211320
LMArena Multi-Turn12721342
WildBench80.2%—

Frequently asked questions

Is Qwen2.5 72B Instruct better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 31.9 on the Noometry Index.

Which is cheaper, Qwen2.5 72B Instruct or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Qwen2.5 72B Instruct lists at $1.40 and $5.60.

Is Qwen2.5 72B Instruct or Trinity Large Thinking better for coding?

They score almost the same on coding (33.2 vs 34.1); test both on your own repository before choosing.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 131K.

How many benchmarks do Qwen2.5 72B Instruct and Trinity Large Thinking share?

17 benchmarks have published results for both models. Qwen2.5 72B Instruct has 43 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper