StepFun, open weights

Step 3

Step 3 by StepFun ranks 149th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.5. Its strongest category is multimodal, where it ranks 86th.

Last verified

Specifications

Noometry rank
#149 of 354
Index score
40.5
Evidence
Confirmed 17 results
Provider
StepFun
Released
Unknown
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
7 tokens/s Kagi
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Step 3 category scores
  1. Coding 40.1
  2. Reasoning 28.4
  3. Math 37.6
  4. Knowledge 36.8
  5. Multimodal 35.5
  6. Multilingual 46.3
  7. Instruction Following 70.4
  8. Long Context 40.3
  9. Writing & Preference 54.3
Step 3 category ranks
CategoryScoreRankResults
Coding40.1#1471
Reasoning28.4#1052
Math37.6#1481
Knowledge36.8#1641
Multimodal35.5#861
Multilingual46.3#1591
Instruction Following70.4#1641
Long Context40.3#1571
Writing & Preference54.3#1513

Strengths and weaknesses

Categories where Step 3 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Step 3: strongest categories
CategoryScorevs medianRank
Reasoning28.4+4.8#105 of 350, top 30%
Coding40.1+1.4#147 of 340, top 44%
Math37.6+1.1#148 of 327, top 46%

Weakest categories

Step 3: weakest categories
CategoryScorevs medianRank
Multimodal35.5−3.0#86 of 128, top 68%
Instruction Following70.4−0.9#164 of 305, top 54%
Multilingual46.3−1.1#159 of 297, top 54%

Closest competitors

The models ranked just above and below Step 3. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Step 3
ModelRankScoreBlended $/MSpeed
Claude Sonnet 4#14540.8$631Compare
Qwen2.5-Max#14640.7——Compare
Nemotron 3 Nano 30B A3B#14740.6$0.0875—Compare
Granite 4.2 8B#14840.5$0.11—Compare
MiniMax M1#15040.3$0.96—Compare
Nvidia Llama 3.3 Nemotron Super 49b v1.5#15140.3$0.40—Compare
Mistral Medium 3.5#15240.2$350Compare
Nemotron 3 Super#15340.1$0.17—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Step 3 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1367#161 of 294, top 55%LMArena2026-10-08

Reasoning

Step 3 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Kagi LLM Benchmark62.3%#38 of 99, top 39%Kagi LLM Benchmark
LMArena Hard Prompts1355#158 of 297, top 54%LMArena2026-10-08

Math

Step 3 Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1366#153 of 285, top 54%LMArena2026-10-08

Knowledge

Step 3 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1333#164 of 273, top 61%LMArena2026-10-08

Multimodal

Step 3 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1177#89 of 122, top 73%LMArena2026-10-09

Multilingual

Step 3 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1327#159 of 297, top 54%LMArena2026-10-08
LMArena Chinese1397#139 of 285, top 49%LMArena2026-10-08
LMArena German1371#115 of 231, top 50%LMArena2026-10-08
LMArena Korean1269#142 of 213, top 67%LMArena2026-10-08
LMArena Russian1331#160 of 283, top 57%LMArena2026-10-08
LMArena Spanish1371#132 of 226, top 59%LMArena2026-10-08

Instruction Following

Step 3 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1332#157 of 298, top 53%LMArena2026-10-08

Long Context

Step 3 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1326#168 of 291, top 58%LMArena2026-10-08

Writing & Preference

Step 3 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1350#159 of 297, top 54%LMArena2026-10-08
LMArena Creative Writing1321#151 of 295, top 52%LMArena2026-10-08
LMArena Multi-Turn1341#160 of 295, top 55%LMArena2026-10-08

Compare Step 3

Other StepFun models

Frequently asked questions

How good is Step 3?

Step 3 by StepFun ranks 149th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.5. Its strongest category is multimodal, where it ranks 86th.

Is Step 3 open source?

Yes. Step 3's weights are downloadable; check the license for commercial terms.

How fast is Step 3?

Step 3 generated about 7 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Step 3's strengths and weaknesses?

Relative to other ranked models, Step 3 places best in reasoning, coding, math and lowest in multimodal, instruction following, multilingual.

What is Step 3 best at?

Its best category is multimodal, where it ranks 86th on Noometry.