Model comparison

Qwen1.5 4b Chat vs Step 3.7 Flash

Step 3.7 Flash is the stronger model overall, scoring 37.3 to 28.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Step 3.7 Flash StepFun

37.3

Rank #207 Reported

Summary

  • The widest gap is in math, where Step 3.7 Flash leads 42.9 to 30.4.

Side by side

Qwen1.5 4b Chat and Step 3.7 Flash specifications
Qwen1.5 4b ChatStep 3.7 Flash
ProviderAlibaba (Qwen)StepFun
Noometry Index28.837.3
Released—2026-05-29
WeightsOpenOpen
Context window—256K
Max output—256K
Input $ / M tokens—$0.18
Output $ / M tokens—$1.11
Results tracked135

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.7 Flash leads

Qwen1.5 4b Chat: 29.1 (#308), Step 3.7 Flash: 40.0 (#150)

Coding benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
SciCode—40%
LMArena Coding999—
ALE-Bench—694.12

Reasoning Step 3.7 Flash leads

Qwen1.5 4b Chat: 18.5 (#279), Step 3.7 Flash: 21.6 (#219)

Reasoning benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
NYT Connections (extended)—39.7%
CritPt—2.3%
LMArena Hard Prompts976—

Math Step 3.7 Flash leads

Qwen1.5 4b Chat: 30.4 (#234), Step 3.7 Flash: 42.9 (#82)

Math benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
MathArena Final-Answer Competitions—68.5%
LMArena Math1026—

Knowledge Not comparable

Qwen1.5 4b Chat: 26.7 (#255), Step 3.7 Flash: —

Knowledge benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
LMArena Expert980—

Multilingual Not comparable

Qwen1.5 4b Chat: 24.1 (#290), Step 3.7 Flash: —

Multilingual benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
LMArena Non-English979—
LMArena Chinese1024—
LMArena German902—
LMArena Russian952—

Instruction Following Not comparable

Qwen1.5 4b Chat: 49.0 (#300), Step 3.7 Flash: —

Instruction Following benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
LMArena Instruction Following978—

Long Context Not comparable

Qwen1.5 4b Chat: 30.1 (#290), Step 3.7 Flash: —

Long Context benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
LMArena Longer Query988—

Writing & Preference Not comparable

Qwen1.5 4b Chat: 23.8 (#309), Step 3.7 Flash: —

Writing & Preference benchmarks
BenchmarkQwen1.5 4b ChatStep 3.7 Flash
LMArena Text997—
LMArena Creative Writing969—
LMArena Multi-Turn977—

Frequently asked questions

Is Qwen1.5 4b Chat better than Step 3.7 Flash?

Step 3.7 Flash is the stronger model overall, scoring 37.3 to 28.8 on the Noometry Index.

Is Qwen1.5 4b Chat or Step 3.7 Flash better for coding?

Step 3.7 Flash scores higher on coding benchmarks: 40.0 versus 29.1 in the Noometry coding category.

How many benchmarks do Qwen1.5 4b Chat and Step 3.7 Flash share?

0 benchmarks have published results for both models. Qwen1.5 4b Chat has 13 scored results on Noometry and Step 3.7 Flash has 5.

Related comparisons

Go deeper