Model comparison

o1 vs Step 1o Turbo 202506

o1 is the stronger model overall, scoring 40.9 to 39.7 on the Noometry Index.

Last verified . 14 shared benchmarks.

o1 OpenAI

40.9

Rank #143 Confirmed

Step 1o Turbo 202506 StepFun

39.7

Rank #160 Confirmed

Summary

  • They share 14 benchmarks with published results for both. o1 scores higher in 7 categories and Step 1o Turbo 202506 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o1 leads 50.3 to 41.0.

Side by side

o1 and Step 1o Turbo 202506 specifications
o1Step 1o Turbo 202506
ProviderOpenAIStepFun
Noometry Index40.939.7
Released2024-09-12—
WeightsProprietaryProprietary
Context window200K—
Max output100K—
Input $ / M tokens$15—
Output $ / M tokens$60—
Results tracked5214

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

o1: 46.1 (#70), Step 1o Turbo 202506: 39.2 (#160)

Coding benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Coding13671339
Aider Polyglot61.7%—
WeirdML47.6%—
LiveBench Coding69.7%—
CadEval56%—
HumanEval+89%—
MBPP+80.2%—

Agentic & Tool Use Not comparable

o1: 24.6 (#117), Step 1o Turbo 202506: —

Agentic & Tool Use benchmarks
Benchmarko1Step 1o Turbo 202506
Cybench10%—
METR Time Horizons51.1%—

Reasoning o1 leads

o1: 27.9 (#111), Step 1o Turbo 202506: 26.8 (#129)

Reasoning benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Hard Prompts13711335
SimpleBench41.7%—
ARC-AGI-130.7%—
Chess Puzzles15%—
EnigmaEval5.7%—
LiveBench Reasoning91.6%—
DTBench74.7%—
LiveBench Data Analysis65.5%—
LMCA22.3%—
Epoch Capabilities Index141.91—
LiveBench75.7%—

Math Too close to call

o1: 36.1 (#175), Step 1o Turbo 202506: 36.6 (#164)

Math benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Math13881318
FrontierMath (Tiers 1-3)14.7%—
OTIS Mock AIME 2024-202573.3%—
LiveBench Math80.3%—
MATH Level 594.7%—
FrontierMath (Feb 2025 set)9.3%—

Knowledge o1 leads

o1: 41.5 (#110), Step 1o Turbo 202506: 36.1 (#176)

Knowledge benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Expert13611308
GPQA Diamond76.8%—
Humanity's Last Exam8%—
SimpleQA Verified41.1%—
Confabulations11.7%—

Multimodal Step 1o Turbo 202506 leads

o1: 34.2 (#93), Step 1o Turbo 202506: 36.1 (#80)

Multimodal benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Vision11681186
GeoBench80%—
VPCT37%—
SpatialViz-Bench41.4%—

Multilingual o1 leads

o1: 48.6 (#142), Step 1o Turbo 202506: 45.3 (#173)

Multilingual benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Non-English13581313
LMArena Chinese13941380
LMArena German13371314
LMArena Russian13561327
LMArena French1344—
LMArena Japanese1346—
LMArena Korean1396—
LMArena Spanish1345—

Instruction Following o1 leads

o1: 74.8 (#86), Step 1o Turbo 202506: 69.2 (#176)

Instruction Following benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Instruction Following13671310
LiveBench Instruction Following81.5%—

Long Context o1 leads

o1: 50.3 (#9), Step 1o Turbo 202506: 41.0 (#146)

Long Context benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Longer Query13781348
Fiction.LiveBench83.3%—

Writing & Preference o1 leads

o1: 55.6 (#144), Step 1o Turbo 202506: 53.1 (#160)

Writing & Preference benchmarks
Benchmarko1Step 1o Turbo 202506
LMArena Text13661336
LMArena Creative Writing13481307
LMArena Multi-Turn13691340
Short-Story Creative Writing70.2%—
LiveBench Language65.4%—

Frequently asked questions

Is o1 better than Step 1o Turbo 202506?

o1 is the stronger model overall, scoring 40.9 to 39.7 on the Noometry Index.

Is o1 or Step 1o Turbo 202506 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 39.2 in the Noometry coding category.

How many benchmarks do o1 and Step 1o Turbo 202506 share?

14 benchmarks have published results for both models. o1 has 52 scored results on Noometry and Step 1o Turbo 202506 has 14.

Related comparisons

Go deeper