Model comparison

Grok Build 0.1 vs Hy4 preview

Hy4 preview is the stronger model overall, scoring 45.3 to 36.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Hy4 preview Tencent

45.3

Rank #73 Reported

Summary

  • The widest gap is in coding, where Hy4 preview leads 51.6 to 43.1.
  • Hy4 preview is cheaper at $0.75 / $2.25 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Hy4 preview accepts more context: 1.05M tokens versus 256K.
  • Hy4 preview has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Hy4 preview specifications
Grok Build 0.1Hy4 preview
ProviderxAITencent
Noometry Index36.445.3
Released2026-04-162026-08-28
WeightsProprietaryOpen
Context window256K1.05M
Max output256K64K
Input $ / M tokens$1$0.75
Output $ / M tokens$2$2.25
Results tracked33

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Grok Build 0.1: 43.1 (#91), Hy4 preview: 51.6 (#38)

Coding benchmarks
BenchmarkGrok Build 0.1Hy4 preview
LMArena WebDev—1632
SciCode50.2%—

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Hy4 preview: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Hy4 preview
GBAEval2.4%—

Reasoning Too close to call

Grok Build 0.1: 32.2 (#77), Hy4 preview: 31.9 (#79)

Reasoning benchmarks
BenchmarkGrok Build 0.1Hy4 preview
NYT Connections (extended)—68.2%
CritPt9.1%—

Math Not comparable

Grok Build 0.1: —, Hy4 preview: 55.7 (#42)

Math benchmarks
BenchmarkGrok Build 0.1Hy4 preview
ProofBench—75%

Frequently asked questions

Is Grok Build 0.1 better than Hy4 preview?

Hy4 preview is the stronger model overall, scoring 45.3 to 36.4 on the Noometry Index.

Which is cheaper, Grok Build 0.1 or Hy4 preview?

Hy4 preview is cheaper. It lists at $0.75 per million input tokens and $2.25 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Hy4 preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 43.1 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 256K.

How many benchmarks do Grok Build 0.1 and Hy4 preview share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Hy4 preview has 3.

Related comparisons

Go deeper