Model comparison

Grok Build 0.1 vs Step 3.5 Flash

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 36.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 22.2.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Step 3.5 Flash specifications
Grok Build 0.1Step 3.5 Flash
ProviderxAIStepFun
Noometry Index36.442.3
Released2026-04-162026-01-29
WeightsProprietaryOpen
Context window256K256K
Max output256K256K
Input $ / M tokens$1$0.10
Output $ / M tokens$2$0.30
Results tracked319

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok Build 0.1: 43.1 (#91), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
SciCode50.2%—
LMArena Coding—1436

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Step 3.5 Flash: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
NYT Connections (extended)—28.4%
CritPt9.1%—
LMArena Hard Prompts—1411

Math Not comparable

Grok Build 0.1: —, Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
MathArena Final-Answer Competitions—66.8%
LMArena Math—1408

Knowledge Not comparable

Grok Build 0.1: —, Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
LMArena Expert—1421

Multilingual Not comparable

Grok Build 0.1: —, Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
LMArena Non-English—1385
LMArena Chinese—1447
LMArena French—1421
LMArena German—1405
LMArena Japanese—1354
LMArena Korean—1352
LMArena Russian—1385
LMArena Spanish—1419

Instruction Following Not comparable

Grok Build 0.1: —, Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
LMArena Instruction Following—1385

Long Context Not comparable

Grok Build 0.1: —, Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
LMArena Longer Query—1402

Writing & Preference Not comparable

Grok Build 0.1: —, Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Step 3.5 Flash
LMArena Text—1403
LMArena Creative Writing—1357
LMArena Multi-Turn—1405

Frequently asked questions

Is Grok Build 0.1 better than Step 3.5 Flash?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 36.4 on the Noometry Index.

Which is cheaper, Grok Build 0.1 or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Step 3.5 Flash better for coding?

They score almost the same on coding (43.1 vs 42.4); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 256K tokens.

How many benchmarks do Grok Build 0.1 and Step 3.5 Flash share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper