Model comparison

Grok-2 (Dec 2024) vs Step 3.5 Flash

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 33.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok-2 (Dec 2024) scores higher in 0 categories and Step 3.5 Flash in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Step 3.5 Flash leads 42.6 to 20.8.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

Grok-2 (Dec 2024) and Step 3.5 Flash specifications
Grok-2 (Dec 2024)Step 3.5 Flash
ProviderxAIStepFun
Noometry Index33.742.3
Released2024-08-132026-01-29
WeightsProprietaryOpen
Context window—256K
Max output—256K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.30
Results tracked3419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.5 Flash leads

Grok-2 (Dec 2024): 33.3 (#258), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Coding12871436
WeirdML22.2%—
LiveBench Coding46.4%—

Reasoning Step 3.5 Flash leads

Grok-2 (Dec 2024): 16.9 (#299), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Hard Prompts12721411
SimpleBench22.7%—
NYT Connections (extended)—28.4%
LiveBench Reasoning54.8%—
DTBench65.2%—
LiveBench Data Analysis54.5%—
Epoch Capabilities Index130.48—
LiveBench54.3%—

Math Step 3.5 Flash leads

Grok-2 (Dec 2024): 20.8 (#284), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Math12831408
MathArena Final-Answer Competitions—66.8%
OTIS Mock AIME 2024-202511.5%—
LiveBench Math54.9%—
MATH Level 563.5%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Step 3.5 Flash leads

Grok-2 (Dec 2024): 29.8 (#233), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Expert12541421
GPQA Diamond53.8%—
Confabulations20.1%—

Multilingual Step 3.5 Flash leads

Grok-2 (Dec 2024): 43.1 (#188), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Non-English12821385
LMArena Chinese12891447
LMArena French13181421
LMArena German12871405
LMArena Japanese12441354
LMArena Korean12371352
LMArena Russian12861385
LMArena Spanish12811419

Instruction Following Step 3.5 Flash leads

Grok-2 (Dec 2024): 66.9 (#202), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Instruction Following12701385
LiveBench Instruction Following69.6%—

Long Context Step 3.5 Flash leads

Grok-2 (Dec 2024): 38.8 (#190), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Longer Query12761402

Writing & Preference Step 3.5 Flash leads

Grok-2 (Dec 2024): 48.6 (#198), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkGrok-2 (Dec 2024)Step 3.5 Flash
LMArena Text13051403
LMArena Creative Writing12841357
LMArena Multi-Turn12901405
Short-Story Creative Writing63.6%—
LiveBench Language45.6%—

Frequently asked questions

Is Grok-2 (Dec 2024) better than Step 3.5 Flash?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 33.7 on the Noometry Index.

Is Grok-2 (Dec 2024) or Step 3.5 Flash better for coding?

Step 3.5 Flash scores higher on coding benchmarks: 42.4 versus 33.3 in the Noometry coding category.

How many benchmarks do Grok-2 (Dec 2024) and Step 3.5 Flash share?

17 benchmarks have published results for both models. Grok-2 (Dec 2024) has 34 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper