Model comparison

GPT-4.1 nano vs Step 3

Step 3 is the stronger model overall, scoring 40.5 to 27.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

GPT-4.1 nano OpenAI

27.9

Rank #327 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 15 benchmarks with published results for both. GPT-4.1 nano scores higher in 0 categories and Step 3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Step 3 leads 28.4 to 8.5.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 33.3% for GPT-4.1 nano and 62.3% for Step 3.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 nano and Step 3 specifications
GPT-4.1 nanoStep 3
ProviderOpenAIStepFun
Noometry Index27.940.5
Released2025-04-14—
WeightsProprietaryOpen
Context window1.05M—
Max output33K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.40—
Results tracked3817

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

GPT-4.1 nano: 24.1 (#330), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Coding13061367
Aider Polyglot8.9%—
SciCode25.9%—
WeirdML19%—

Agentic & Tool Use Not comparable

GPT-4.1 nano: 26.5 (#104), Step 3: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 nanoStep 3
Berkeley Function Calling Leaderboard33%—

Reasoning Step 3 leads

GPT-4.1 nano: 8.5 (#349), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkGPT-4.1 nanoStep 3
Kagi LLM Benchmark33.3%62.3%
LMArena Hard Prompts12861355
ARC-AGI-20%—
ARC-AGI-10%—
CritPt0%—
DTBench52.5%—
LMCA5.5%—
Epoch Capabilities Index129.62—

Math Step 3 leads

GPT-4.1 nano: 26.9 (#252), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Math12741366
OTIS Mock AIME 2024-202528.9%—
Omni-MATH36.7%—
MATH Level 570%—
FrontierMath (Feb 2025 set)1%—

Knowledge Step 3 leads

GPT-4.1 nano: 21.8 (#273), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Expert12721333
GPQA Diamond48.9%—
SimpleQA Verified6%—
MMLU-Pro55%—
GPQA (HELM)50.7%—

Multimodal Step 3 leads

GPT-4.1 nano: 29.2 (#113), Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Vision10631177

Multilingual Step 3 leads

GPT-4.1 nano: 41.6 (#205), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Non-English12601327
LMArena Chinese12701397
LMArena German12881371
LMArena Russian12611331
LMArena Japanese1198—
LMArena Korean—1269
LMArena Spanish—1371

Instruction Following Step 3 leads

GPT-4.1 nano: 67.8 (#193), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Instruction Following12671332
IFEval84.3%—

Long Context Step 3 leads

GPT-4.1 nano: 23.7 (#296), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Longer Query12831326
Fiction.LiveBench25%—

Writing & Preference Step 3 leads

GPT-4.1 nano: 40.5 (#243), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkGPT-4.1 nanoStep 3
LMArena Text12851350
LMArena Creative Writing12601321
LMArena Multi-Turn12771341
EQ-Bench Creative Writing946—
WildBench81.2%—

Frequently asked questions

Is GPT-4.1 nano better than Step 3?

Step 3 is the stronger model overall, scoring 40.5 to 27.9 on the Noometry Index.

Is GPT-4.1 nano or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 24.1 in the Noometry coding category.

How many benchmarks do GPT-4.1 nano and Step 3 share?

15 benchmarks have published results for both models. GPT-4.1 nano has 38 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper