Model comparison

Llama2 70b Steerlm Chat vs Step 3.7 Flash

Step 3.7 Flash is the stronger model overall, scoring 37.3 to 31.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Step 3.7 Flash StepFun

37.3

Rank #207 Reported

Summary

  • The widest gap is in math, where Step 3.7 Flash leads 42.9 to 31.3.

Side by side

Llama2 70b Steerlm Chat and Step 3.7 Flash specifications
Llama2 70b Steerlm ChatStep 3.7 Flash
ProviderNVIDIAStepFun
Noometry Index31.837.3
Released—2026-05-29
WeightsOpenOpen
Context window—256K
Max output—256K
Input $ / M tokens—$0.18
Output $ / M tokens—$1.11
Results tracked95

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.7 Flash leads

Llama2 70b Steerlm Chat: 29.9 (#300), Step 3.7 Flash: 40.0 (#150)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
SciCode—40%
LMArena Coding1025—
ALE-Bench—694.12

Reasoning Step 3.7 Flash leads

Llama2 70b Steerlm Chat: 20.0 (#246), Step 3.7 Flash: 21.6 (#219)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
NYT Connections (extended)—39.7%
CritPt—2.3%
LMArena Hard Prompts1047—

Math Step 3.7 Flash leads

Llama2 70b Steerlm Chat: 31.3 (#226), Step 3.7 Flash: 42.9 (#82)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
MathArena Final-Answer Competitions—68.5%
LMArena Math1072—

Multilingual Not comparable

Llama2 70b Steerlm Chat: 28.8 (#270), Step 3.7 Flash: —

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
LMArena Non-English1063—

Instruction Following Not comparable

Llama2 70b Steerlm Chat: 54.2 (#279), Step 3.7 Flash: —

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
LMArena Instruction Following1060—

Long Context Not comparable

Llama2 70b Steerlm Chat: 30.4 (#288), Step 3.7 Flash: —

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
LMArena Longer Query998—

Writing & Preference Not comparable

Llama2 70b Steerlm Chat: 31.6 (#283), Step 3.7 Flash: —

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatStep 3.7 Flash
LMArena Text1098—
LMArena Creative Writing1091—
LMArena Multi-Turn1058—

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Step 3.7 Flash?

Step 3.7 Flash is the stronger model overall, scoring 37.3 to 31.8 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Step 3.7 Flash better for coding?

Step 3.7 Flash scores higher on coding benchmarks: 40.0 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Step 3.7 Flash share?

0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Step 3.7 Flash has 5.

Related comparisons

Go deeper