Model comparison

Step 1o Turbo 202506 vs Trinity Large Thinking

Step 1o Turbo 202506 is the stronger model overall, scoring 39.7 to 38.6 on the Noometry Index.

Last verified . 13 shared benchmarks.

Step 1o Turbo 202506 StepFun

39.7

Rank #160 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Step 1o Turbo 202506 scores higher in 2 categories and Trinity Large Thinking in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Step 1o Turbo 202506 leads 26.8 to 16.9.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Step 1o Turbo 202506 and Trinity Large Thinking specifications
Step 1o Turbo 202506Trinity Large Thinking
ProviderStepFunArcee AI
Noometry Index39.738.6
Released—2026-04-01
WeightsProprietaryOpen
Context window—262K
Max output—80K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.80
Results tracked1424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 1o Turbo 202506 leads

Step 1o Turbo 202506: 39.2 (#160), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Coding13391381
LMArena WebDev—1238
SciCode—36.1%

Reasoning Step 1o Turbo 202506 leads

Step 1o Turbo 202506: 26.8 (#129), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Hard Prompts13351350
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

Step 1o Turbo 202506: 36.6 (#164), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Math13181366

Knowledge Trinity Large Thinking leads

Step 1o Turbo 202506: 36.1 (#176), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Expert13081360
Vectara Hallucination Rate—6.9%

Multimodal Not comparable

Step 1o Turbo 202506: 36.1 (#80), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Vision1186—

Multilingual Too close to call

Step 1o Turbo 202506: 45.3 (#173), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Non-English13131325
LMArena Chinese13801373
LMArena German13141356
LMArena Russian13271337
LMArena French—1374
LMArena Japanese—1311
LMArena Korean—1306
LMArena Spanish—1357

Instruction Following Trinity Large Thinking leads

Step 1o Turbo 202506: 69.2 (#176), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Instruction Following13101334

Long Context Too close to call

Step 1o Turbo 202506: 41.0 (#146), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Longer Query13481355

Writing & Preference Too close to call

Step 1o Turbo 202506: 53.1 (#160), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkStep 1o Turbo 202506Trinity Large Thinking
LMArena Text13361340
LMArena Creative Writing13071320
LMArena Multi-Turn13401342

Frequently asked questions

Is Step 1o Turbo 202506 better than Trinity Large Thinking?

Step 1o Turbo 202506 is the stronger model overall, scoring 39.7 to 38.6 on the Noometry Index.

Is Step 1o Turbo 202506 or Trinity Large Thinking better for coding?

Step 1o Turbo 202506 scores higher on coding benchmarks: 39.2 versus 34.1 in the Noometry coding category.

How many benchmarks do Step 1o Turbo 202506 and Trinity Large Thinking share?

13 benchmarks have published results for both models. Step 1o Turbo 202506 has 14 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper