Model comparison
Llama2 70b Steerlm Chat vs Step 3.7 Flash
Step 3.7 Flash is the stronger model overall, scoring 37.3 to 31.8 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in math, where Step 3.7 Flash leads 42.9 to 31.3.
Side by side
| Llama2 70b Steerlm Chat | Step 3.7 Flash | |
|---|---|---|
| Provider | NVIDIA | StepFun |
| Noometry Index | 31.8 | 37.3 |
| Released | — | 2026-05-29 |
| Weights | Open | Open |
| Context window | — | 256K |
| Max output | — | 256K |
| Input $ / M tokens | — | $0.18 |
| Output $ / M tokens | — | $1.11 |
| Results tracked | 9 | 5 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Step 3.7 Flash leads
Llama2 70b Steerlm Chat: 29.9 (#300), Step 3.7 Flash: 40.0 (#150)
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| SciCode | — | 40% |
| LMArena Coding | 1025 | — |
| ALE-Bench | — | 694.12 |
Reasoning Step 3.7 Flash leads
Llama2 70b Steerlm Chat: 20.0 (#246), Step 3.7 Flash: 21.6 (#219)
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| NYT Connections (extended) | — | 39.7% |
| CritPt | — | 2.3% |
| LMArena Hard Prompts | 1047 | — |
Math Step 3.7 Flash leads
Llama2 70b Steerlm Chat: 31.3 (#226), Step 3.7 Flash: 42.9 (#82)
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| MathArena Final-Answer Competitions | — | 68.5% |
| LMArena Math | 1072 | — |
Multilingual Not comparable
Llama2 70b Steerlm Chat: 28.8 (#270), Step 3.7 Flash: —
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Non-English | 1063 | — |
Instruction Following Not comparable
Llama2 70b Steerlm Chat: 54.2 (#279), Step 3.7 Flash: —
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Instruction Following | 1060 | — |
Long Context Not comparable
Llama2 70b Steerlm Chat: 30.4 (#288), Step 3.7 Flash: —
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Longer Query | 998 | — |
Writing & Preference Not comparable
Llama2 70b Steerlm Chat: 31.6 (#283), Step 3.7 Flash: —
| Benchmark | Llama2 70b Steerlm Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Text | 1098 | — |
| LMArena Creative Writing | 1091 | — |
| LMArena Multi-Turn | 1058 | — |
Frequently asked questions
Is Llama2 70b Steerlm Chat better than Step 3.7 Flash?
Step 3.7 Flash is the stronger model overall, scoring 37.3 to 31.8 on the Noometry Index.
Is Llama2 70b Steerlm Chat or Step 3.7 Flash better for coding?
Step 3.7 Flash scores higher on coding benchmarks: 40.0 versus 29.9 in the Noometry coding category.
How many benchmarks do Llama2 70b Steerlm Chat and Step 3.7 Flash share?
0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Step 3.7 Flash has 5.