Model comparison

GPT-4 Turbo vs Step 3

Step 3 is the stronger model overall, scoring 40.5 to 30.5 on the Noometry Index.

Last verified . 16 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 16 benchmarks with published results for both. GPT-4 Turbo scores higher in 0 categories and Step 3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Step 3 leads 37.6 to 9.0.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

GPT-4 Turbo and Step 3 specifications
GPT-4 TurboStep 3
ProviderOpenAIStepFun
Noometry Index30.540.5
Released2023-11-06—
WeightsProprietaryOpen
Context window128K—
Max output4K—
Input $ / M tokens$10—
Output $ / M tokens$30—
Results tracked3617

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

GPT-4 Turbo: 33.8 (#249), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Coding12681367
WeirdML18%—
BigCodeBench Instruct48.2%—
BigCodeBench Complete58.2%—
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Step 3: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboStep 3
METR Time Horizons36.7%—

Reasoning Step 3 leads

GPT-4 Turbo: 15.3 (#317), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Hard Prompts12511355
SimpleBench25.1%—
Kagi LLM Benchmark—62.3%
Chess Puzzles6%—
DTBench61.6%—
LMCA9.8%—
Epoch Capabilities Index127.25—
ForecastBench59.4—

Math Step 3 leads

GPT-4 Turbo: 9.0 (#322), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Math12721366
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.7%—
MATH Level 546.7%—

Knowledge Step 3 leads

GPT-4 Turbo: 24.3 (#268), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Expert12231333
GPQA Diamond46.6%—
Confabulations28.4%—
MMLU81.3%—

Multimodal Step 3 leads

GPT-4 Turbo: 30.6 (#110), Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Vision10901177

Multilingual Step 3 leads

GPT-4 Turbo: 40.5 (#216), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Non-English12451327
LMArena Chinese12421397
LMArena German12591371
LMArena Korean11871269
LMArena Russian12591331
LMArena Spanish12601371
LMArena French1276—
LMArena Japanese1194—

Instruction Following Step 3 leads

GPT-4 Turbo: 65.8 (#216), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Instruction Following12491332

Long Context Step 3 leads

GPT-4 Turbo: 38.0 (#206), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Longer Query12541326

Writing & Preference Step 3 leads

GPT-4 Turbo: 47.7 (#206), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkGPT-4 TurboStep 3
LMArena Text12721350
LMArena Creative Writing12691321
LMArena Multi-Turn12671341

Frequently asked questions

Is GPT-4 Turbo better than Step 3?

Step 3 is the stronger model overall, scoring 40.5 to 30.5 on the Noometry Index.

Is GPT-4 Turbo or Step 3 better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 33.8 in the Noometry coding category.

How many benchmarks do GPT-4 Turbo and Step 3 share?

16 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper