Model comparison

Step 3 vs Yi-34B

Step 3 is the stronger model overall, scoring 40.5 to 27.8 on the Noometry Index.

Last verified . 15 shared benchmarks.

Step 3 StepFun

40.5

Rank #149 Confirmed

Yi-34B 01.AI

27.8

Rank #329 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Step 3 scores higher in 8 categories and Yi-34B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Step 3 leads 36.8 to 7.5.

Side by side

Step 3 and Yi-34B specifications
Step 3Yi-34B
ProviderStepFun01.AI
Noometry Index40.527.8
Released—2023-11-02
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1723

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3 leads

Step 3: 40.1 (#147), Yi-34B: 32.3 (#274)

Coding benchmarks
BenchmarkStep 3Yi-34B
LMArena Coding13671112

Reasoning Step 3 leads

Step 3: 28.4 (#105), Yi-34B: 21.2 (#226)

Reasoning benchmarks
BenchmarkStep 3Yi-34B
LMArena Hard Prompts13551104
Kagi LLM Benchmark62.3%—
BIG-Bench Hard—71.7%
Epoch Capabilities Index—117.39

Math Step 3 leads

Step 3: 37.6 (#148), Yi-34B: 21.6 (#282)

Math benchmarks
BenchmarkStep 3Yi-34B
LMArena Math13661114
MATH Level 5—5.1%
GSM8K—76%

Knowledge Step 3 leads

Step 3: 36.8 (#164), Yi-34B: 7.5 (#309)

Knowledge benchmarks
BenchmarkStep 3Yi-34B
LMArena Expert13331061
GPQA Diamond—14.7%
MMLU—76.3%

Multimodal Not comparable

Step 3: 35.5 (#86), Yi-34B: —

Multimodal benchmarks
BenchmarkStep 3Yi-34B
LMArena Vision1177—

Multilingual Step 3 leads

Step 3: 46.3 (#159), Yi-34B: 29.7 (#264)

Multilingual benchmarks
BenchmarkStep 3Yi-34B
LMArena Non-English13271079
LMArena Chinese13971176
LMArena German13711042
LMArena Korean1269959
LMArena Russian13311050
LMArena Spanish13711070
LMArena French—1081
LMArena Japanese—993

Instruction Following Step 3 leads

Step 3: 70.4 (#164), Yi-34B: 56.2 (#274)

Instruction Following benchmarks
BenchmarkStep 3Yi-34B
LMArena Instruction Following13321091

Long Context Step 3 leads

Step 3: 40.3 (#157), Yi-34B: 33.2 (#264)

Long Context benchmarks
BenchmarkStep 3Yi-34B
LMArena Longer Query13261094

Writing & Preference Step 3 leads

Step 3: 54.3 (#151), Yi-34B: 34.1 (#273)

Writing & Preference benchmarks
BenchmarkStep 3Yi-34B
LMArena Text13501129
LMArena Creative Writing13211108
LMArena Multi-Turn13411113

Frequently asked questions

Is Step 3 better than Yi-34B?

Step 3 is the stronger model overall, scoring 40.5 to 27.8 on the Noometry Index.

Is Step 3 or Yi-34B better for coding?

Step 3 scores higher on coding benchmarks: 40.1 versus 32.3 in the Noometry coding category.

How many benchmarks do Step 3 and Yi-34B share?

15 benchmarks have published results for both models. Step 3 has 17 scored results on Noometry and Yi-34B has 23.

Related comparisons

Go deeper