Model comparison
Qwen1.5 4b Chat vs Step 3.7 Flash
Step 3.7 Flash is the stronger model overall, scoring 37.3 to 28.8 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in math, where Step 3.7 Flash leads 42.9 to 30.4.
Side by side
| Qwen1.5 4b Chat | Step 3.7 Flash | |
|---|---|---|
| Provider | Alibaba (Qwen) | StepFun |
| Noometry Index | 28.8 | 37.3 |
| Released | — | 2026-05-29 |
| Weights | Open | Open |
| Context window | — | 256K |
| Max output | — | 256K |
| Input $ / M tokens | — | $0.18 |
| Output $ / M tokens | — | $1.11 |
| Results tracked | 13 | 5 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Step 3.7 Flash leads
Qwen1.5 4b Chat: 29.1 (#308), Step 3.7 Flash: 40.0 (#150)
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| SciCode | — | 40% |
| LMArena Coding | 999 | — |
| ALE-Bench | — | 694.12 |
Reasoning Step 3.7 Flash leads
Qwen1.5 4b Chat: 18.5 (#279), Step 3.7 Flash: 21.6 (#219)
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| NYT Connections (extended) | — | 39.7% |
| CritPt | — | 2.3% |
| LMArena Hard Prompts | 976 | — |
Math Step 3.7 Flash leads
Qwen1.5 4b Chat: 30.4 (#234), Step 3.7 Flash: 42.9 (#82)
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| MathArena Final-Answer Competitions | — | 68.5% |
| LMArena Math | 1026 | — |
Knowledge Not comparable
Qwen1.5 4b Chat: 26.7 (#255), Step 3.7 Flash: —
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Expert | 980 | — |
Multilingual Not comparable
Qwen1.5 4b Chat: 24.1 (#290), Step 3.7 Flash: —
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Non-English | 979 | — |
| LMArena Chinese | 1024 | — |
| LMArena German | 902 | — |
| LMArena Russian | 952 | — |
Instruction Following Not comparable
Qwen1.5 4b Chat: 49.0 (#300), Step 3.7 Flash: —
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Instruction Following | 978 | — |
Long Context Not comparable
Qwen1.5 4b Chat: 30.1 (#290), Step 3.7 Flash: —
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Longer Query | 988 | — |
Writing & Preference Not comparable
Qwen1.5 4b Chat: 23.8 (#309), Step 3.7 Flash: —
| Benchmark | Qwen1.5 4b Chat | Step 3.7 Flash |
|---|---|---|
| LMArena Text | 997 | — |
| LMArena Creative Writing | 969 | — |
| LMArena Multi-Turn | 977 | — |
Frequently asked questions
Is Qwen1.5 4b Chat better than Step 3.7 Flash?
Step 3.7 Flash is the stronger model overall, scoring 37.3 to 28.8 on the Noometry Index.
Is Qwen1.5 4b Chat or Step 3.7 Flash better for coding?
Step 3.7 Flash scores higher on coding benchmarks: 40.0 versus 29.1 in the Noometry coding category.
How many benchmarks do Qwen1.5 4b Chat and Step 3.7 Flash share?
0 benchmarks have published results for both models. Qwen1.5 4b Chat has 13 scored results on Noometry and Step 3.7 Flash has 5.