Model comparison

Llama 3.1 Tulu 3 8b vs Step 5 Preview

Step 5 Preview is the stronger model overall, scoring 47.9 to 35.7 on the Noometry Index.

Last verified . 11 shared benchmarks.

Step 5 Preview StepFun

47.9

Rank #58 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.1 Tulu 3 8b scores higher in 0 categories and Step 5 Preview in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Step 5 Preview leads 62.8 to 39.7.
  • Llama 3.1 Tulu 3 8b has downloadable open weights; the other is API-only.

Side by side

Llama 3.1 Tulu 3 8b and Step 5 Preview specifications
Llama 3.1 Tulu 3 8bStep 5 Preview
ProviderAllen Institute for AI (Ai2)StepFun
Noometry Index35.747.9
Released—2026-09-16
WeightsOpenProprietary
Context window—1.02M
Max output—66K
Input $ / M tokens—$1
Output $ / M tokens—$2.70
Results tracked1118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 34.4 (#235), Step 5 Preview: 51.8 (#36)

Coding benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Coding11831480
LMArena WebDev—1564
SciCode—58.9%

Reasoning Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 22.8 (#188), Step 5 Preview: 40.0 (#57)

Reasoning benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Hard Prompts11741465
CritPt—20.9%

Math Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 33.9 (#198), Step 5 Preview: 46.1 (#72)

Math benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Math11951470
ProofBench—42%

Knowledge Not comparable

Llama 3.1 Tulu 3 8b: —, Step 5 Preview: 41.2 (#112)

Knowledge benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Expert—1470

Multimodal Not comparable

Llama 3.1 Tulu 3 8b: —, Step 5 Preview: 41.0 (#41)

Multimodal benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Vision—1267

Multilingual Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 35.4 (#246), Step 5 Preview: 53.8 (#54)

Multilingual benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Non-English11691432
LMArena Chinese11761519
LMArena Russian11931436
LMArena Spanish—1439

Instruction Following Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 61.3 (#246), Step 5 Preview: 76.0 (#50)

Instruction Following benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Instruction Following11741444

Long Context Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 35.8 (#239), Step 5 Preview: 44.5 (#64)

Long Context benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Longer Query11811455

Writing & Preference Step 5 Preview leads

Llama 3.1 Tulu 3 8b: 39.7 (#245), Step 5 Preview: 62.8 (#71)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Tulu 3 8bStep 5 Preview
LMArena Text11931442
LMArena Creative Writing11821410
LMArena Multi-Turn11541447

Frequently asked questions

Is Llama 3.1 Tulu 3 8b better than Step 5 Preview?

Step 5 Preview is the stronger model overall, scoring 47.9 to 35.7 on the Noometry Index.

Is Llama 3.1 Tulu 3 8b or Step 5 Preview better for coding?

Step 5 Preview scores higher on coding benchmarks: 51.8 versus 34.4 in the Noometry coding category.

How many benchmarks do Llama 3.1 Tulu 3 8b and Step 5 Preview share?

11 benchmarks have published results for both models. Llama 3.1 Tulu 3 8b has 11 scored results on Noometry and Step 5 Preview has 18.

Related comparisons

Go deeper