Model comparison

Qwen3.5 Max Preview vs Trinity Large Thinking

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 38.6 on the Noometry Index.

Last verified . 17 shared benchmarks.

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Qwen3.5 Max Preview scores higher in 8 categories and Trinity Large Thinking in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5 Max Preview leads 30.8 to 16.9.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Qwen3.5 Max Preview and Trinity Large Thinking specifications
Qwen3.5 Max PreviewTrinity Large Thinking
ProviderAlibaba (Qwen)Arcee AI
Noometry Index45.338.6
Released—2026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked1724

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 44.0 (#77), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Coding14871381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 30.8 (#84), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Hard Prompts14831350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 40.1 (#94), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Math14741366

Knowledge Too close to call

Qwen3.5 Max Preview: 41.8 (#107), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Expert14891360
Vectara Hallucination Rate—6.9%

Multilingual Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 56.2 (#22), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Non-English14651325
LMArena Chinese15341373
LMArena French14841374
LMArena German14871356
LMArena Japanese14951311
LMArena Korean14381306
LMArena Russian14711337
LMArena Spanish14701357

Instruction Following Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 77.0 (#31), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Instruction Following14671334

Long Context Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 45.2 (#45), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Longer Query14761355

Writing & Preference Qwen3.5 Max Preview leads

Qwen3.5 Max Preview: 66.0 (#41), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkQwen3.5 Max PreviewTrinity Large Thinking
LMArena Text14701340
LMArena Creative Writing14641320
LMArena Multi-Turn14781342

Frequently asked questions

Is Qwen3.5 Max Preview better than Trinity Large Thinking?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 38.6 on the Noometry Index.

Is Qwen3.5 Max Preview or Trinity Large Thinking better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 34.1 in the Noometry coding category.

How many benchmarks do Qwen3.5 Max Preview and Trinity Large Thinking share?

17 benchmarks have published results for both models. Qwen3.5 Max Preview has 17 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper