Model comparison

Phi 3 Small 8k Instruct vs Step 2 16k Exp 202412

Step 2 16k Exp 202412 is the stronger model overall, scoring 39.2 to 29.3 on the Noometry Index.

Last verified . 12 shared benchmarks.

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Step 2 16k Exp 202412 StepFun

39.2

Rank #171 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Phi 3 Small 8k Instruct scores higher in 0 categories and Step 2 16k Exp 202412 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Step 2 16k Exp 202412 leads 52.3 to 31.1.
  • Phi 3 Small 8k Instruct has downloadable open weights; the other is API-only.

Side by side

Phi 3 Small 8k Instruct and Step 2 16k Exp 202412 specifications
Phi 3 Small 8k InstructStep 2 16k Exp 202412
ProviderMicrosoftStepFun
Noometry Index29.339.2
Released2024-04-23—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3212

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 27.9 (#318), Step 2 16k Exp 202412: 38.5 (#174)

Coding benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Coding11011316
LiveBench Coding20.3%—

Reasoning Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 14.8 (#323), Step 2 16k Exp 202412: 25.9 (#142)

Reasoning benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Hard Prompts11001299
LiveBench Reasoning15.9%—
LiveBench Data Analysis30.3%—
Adversarial NLI58.1%—
BIG-Bench Hard79.1%—
HellaSwag77%—
LiveBench24%—
WinoGrande81.5%—

Math Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 27.6 (#248), Step 2 16k Exp 202412: 36.3 (#169)

Math benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Math11511304
LiveBench Math17.6%—

Knowledge Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 29.1 (#240), Step 2 16k Exp 202412: 35.2 (#187)

Knowledge benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Expert10671279
ARC (AI2) Challenge90.7%—
MMLU75.7%—
OpenBookQA88%—
TriviaQA58.1%—

Multilingual Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 28.5 (#272), Step 2 16k Exp 202412: 43.7 (#181)

Multilingual benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Non-English10581290
LMArena Chinese10611331
LMArena Russian11111324
LMArena French1135—
LMArena German1080—
LMArena Japanese966—
LMArena Korean894—
LMArena Spanish1111—

Instruction Following Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 51.9 (#292), Step 2 16k Exp 202412: 67.9 (#192)

Instruction Following benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Instruction Following10871287
LiveBench Instruction Following47.2%—

Long Context Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 33.0 (#267), Step 2 16k Exp 202412: 39.7 (#170)

Long Context benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Longer Query10881307

Writing & Preference Step 2 16k Exp 202412 leads

Phi 3 Small 8k Instruct: 31.1 (#284), Step 2 16k Exp 202412: 52.3 (#173)

Writing & Preference benchmarks
BenchmarkPhi 3 Small 8k InstructStep 2 16k Exp 202412
LMArena Text11101321
LMArena Creative Writing10831328
LMArena Multi-Turn10681292
LiveBench Language12.9%—

Frequently asked questions

Is Phi 3 Small 8k Instruct better than Step 2 16k Exp 202412?

Step 2 16k Exp 202412 is the stronger model overall, scoring 39.2 to 29.3 on the Noometry Index.

Is Phi 3 Small 8k Instruct or Step 2 16k Exp 202412 better for coding?

Step 2 16k Exp 202412 scores higher on coding benchmarks: 38.5 versus 27.9 in the Noometry coding category.

How many benchmarks do Phi 3 Small 8k Instruct and Step 2 16k Exp 202412 share?

12 benchmarks have published results for both models. Phi 3 Small 8k Instruct has 32 scored results on Noometry and Step 2 16k Exp 202412 has 12.

Related comparisons

Go deeper