Model comparison

Chatgpt 4o Latest 20250326 vs Trinity Large Thinking

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 38.6 on the Noometry Index.

Last verified . 17 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 7 categories and Trinity Large Thinking in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Chatgpt 4o Latest 20250326 leads 33.8 to 16.9.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Trinity Large Thinking specifications
Chatgpt 4o Latest 20250326Trinity Large Thinking
ProviderOpenAIArcee AI
Noometry Index43.838.6
Released—2026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Coding14131381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Hard Prompts14241350
Kagi LLM Benchmark75%—
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Too close to call

Chatgpt 4o Latest 20250326: 38.6 (#134), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Math14071366

Knowledge Trinity Large Thinking leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Expert14011360
Confabulations16.6%—
Vectara Hallucination Rate—6.9%

Multimodal Not comparable

Chatgpt 4o Latest 20250326: 39.6 (#58), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Vision1243—

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Non-English14191325
LMArena Chinese14571373
LMArena French14461374
LMArena German14231356
LMArena Japanese14051311
LMArena Korean13961306
LMArena Russian14291337
LMArena Spanish14341357

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Instruction Following14031334

Long Context Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Longer Query14131355

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Trinity Large Thinking
LMArena Text14291340
LMArena Creative Writing14051320
LMArena Multi-Turn14541342
EQ-Bench Creative Writing1501—

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Trinity Large Thinking?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 38.6 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Trinity Large Thinking better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 34.1 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Trinity Large Thinking share?

17 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper