Model comparison

Grok Build 0.1 vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 36.4 on the Noometry Index. Grok Build 0.1 costs 2.1× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Grok Build 0.1 scores higher in 1 category and Inkling in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 34.5.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Grok Build 0.1 accepts more context: 256K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Inkling specifications
Grok Build 0.1Inkling
ProviderxAIThinking Machines Lab
Noometry Index36.444.1
Released2026-04-162026-07-15
WeightsProprietaryOpen
Context window256K66K
Max output256K66K
Input $ / M tokens$1$1.87
Output $ / M tokens$2$4.68
Results tracked341

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGrok Build 0.1Inkling
SciCode50.2%47%
FrontierCode—14%
LMArena WebDev—1413
FrontierSWE—4.1%
WeirdML—32.3%
LMArena Coding—1464
ALE-Bench—946

Agentic & Tool Use Inkling leads

Grok Build 0.1: 22.7 (#129), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Inkling
APEX-Agents—33.8%
τ²-bench Banking—25%
GBAEval2.4%—

Reasoning Inkling leads

Grok Build 0.1: 32.2 (#77), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGrok Build 0.1Inkling
CritPt9.1%5.4%
ARC-AGI-2—36.5%
SimpleBench—50%
ARC-AGI-1—79.5%
Chess Puzzles—21%
LMArena Hard Prompts—1451
DTBench—87.5%
LMCA—37.6%
Epoch Capabilities Index—148.54

Math Not comparable

Grok Build 0.1: —, Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGrok Build 0.1Inkling
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—88.9%
ProofBench—0%
LMArena Math—1479

Knowledge Not comparable

Grok Build 0.1: —, Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGrok Build 0.1Inkling
GPQA Diamond—88.3%
SimpleQA Verified—40.3%
LMArena Expert—1465

Multilingual Not comparable

Grok Build 0.1: —, Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGrok Build 0.1Inkling
LMArena Non-English—1434
LMArena Chinese—1490
LMArena French—1458
LMArena German—1446
LMArena Japanese—1429
LMArena Korean—1404
LMArena Russian—1429
LMArena Spanish—1448

Instruction Following Not comparable

Grok Build 0.1: —, Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Inkling
LMArena Instruction Following—1426

Long Context Not comparable

Grok Build 0.1: —, Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGrok Build 0.1Inkling
LMArena Longer Query—1434

Writing & Preference Not comparable

Grok Build 0.1: —, Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Inkling
LMArena Text—1441
LMArena Creative Writing—1387
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226
LMArena Multi-Turn—1436

Frequently asked questions

Is Grok Build 0.1 better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 36.4 on the Noometry Index. Grok Build 0.1 costs 2.1× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Inkling?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Grok Build 0.1 or Inkling better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 66K.

How many benchmarks do Grok Build 0.1 and Inkling share?

2 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper