Model comparison

Command A vs Trinity Large Thinking

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 36.5 on the Noometry Index.

Last verified . 18 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Command A scores higher in 1 category and Trinity Large Thinking in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Trinity Large Thinking leads 34.1 to 27.2.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $2.50 / $10 for Command A.
  • Trinity Large Thinking accepts more context: 262K tokens versus 256K.

Side by side

Command A and Trinity Large Thinking specifications
Command ATrinity Large Thinking
ProviderCohereArcee AI
Noometry Index36.538.6
Released2025-03-132026-04-01
WeightsOpenOpen
Context window256K262K
Max output8K80K
Input $ / M tokens$2.50$0.25
Output $ / M tokens$10$0.80
Results tracked2424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Trinity Large Thinking leads

Command A: 27.2 (#322), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Coding13301381
Aider Polyglot12%—
LMArena WebDev—1238
SciCode—36.1%

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkCommand ATrinity Large Thinking
Berkeley Function Calling Leaderboard57.1%—

Reasoning Command A leads

Command A: 18.3 (#283), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Hard Prompts13261350
Kagi LLM Benchmark28.8%—
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
DTBench61.3%—
LMCA10.3%—
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Command A: 36.2 (#171), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Math13001366

Knowledge Trinity Large Thinking leads

Command A: 37.1 (#159), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkCommand ATrinity Large Thinking
Vectara Hallucination Rate9.3%6.9%
LMArena Expert12951360

Multilingual Too close to call

Command A: 45.3 (#170), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Non-English13131325
LMArena Chinese13271373
LMArena French13511374
LMArena German13411356
LMArena Japanese12851311
LMArena Korean12851306
LMArena Russian13141337
LMArena Spanish13471357

Instruction Following Trinity Large Thinking leads

Command A: 69.1 (#177), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Instruction Following13091334

Long Context Too close to call

Command A: 40.6 (#151), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Longer Query13341355

Writing & Preference Trinity Large Thinking leads

Command A: 47.6 (#208), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkCommand ATrinity Large Thinking
LMArena Text13311340
LMArena Creative Writing13191320
LMArena Multi-Turn13391342
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than Trinity Large Thinking?

Trinity Large Thinking is the stronger model overall, scoring 38.6 to 36.5 on the Noometry Index.

Which is cheaper, Command A or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or Trinity Large Thinking better for coding?

Trinity Large Thinking scores higher on coding benchmarks: 34.1 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 256K.

How many benchmarks do Command A and Trinity Large Thinking share?

18 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper