Model comparison

Grok Build 0.1 vs o3-pro

o3-pro is the stronger model overall, scoring 42.9 to 36.4 on the Noometry Index. Grok Build 0.1 costs 28× less per token, which makes it the better buy when o3-pro's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • The widest gap is in coding, where o3-pro leads 55.5 to 43.1.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $20 / $80 for o3-pro.
  • Grok Build 0.1 accepts more context: 256K tokens versus 200K.

Side by side

Grok Build 0.1 and o3-pro specifications
Grok Build 0.1o3-pro
ProviderxAIOpenAI
Noometry Index36.442.9
Released2026-04-162025-06-10
WeightsProprietaryProprietary
Context window256K200K
Max output256K100K
Input $ / M tokens$1$20
Output $ / M tokens$2$80
Results tracked312

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-pro leads

Grok Build 0.1: 43.1 (#91), o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkGrok Build 0.1o3-pro
Aider Polyglot—84.9%
SciCode50.2%—
WeirdML—58.2%

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), o3-pro: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1o3-pro
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkGrok Build 0.1o3-pro
ARC-AGI-2—4.9%
Kagi LLM Benchmark—72.1%
ARC-AGI-1—59.3%
CritPt9.1%—
DTBench—86.9%
LMCA—38.5%
Epoch Capabilities Index—147.42

Knowledge Not comparable

Grok Build 0.1: —, o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkGrok Build 0.1o3-pro
Confabulations—14.2%
Vectara Hallucination Rate—23.3%

Long Context Not comparable

Grok Build 0.1: —, o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkGrok Build 0.1o3-pro
Fiction.LiveBench—97.2%

Writing & Preference Not comparable

Grok Build 0.1: —, o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1o3-pro
Short-Story Creative Writing—84.4%

Frequently asked questions

Is Grok Build 0.1 better than o3-pro?

o3-pro is the stronger model overall, scoring 42.9 to 36.4 on the Noometry Index. Grok Build 0.1 costs 28× less per token, which makes it the better buy when o3-pro's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or o3-pro?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; o3-pro lists at $20 and $80.

Is Grok Build 0.1 or o3-pro better for coding?

o3-pro scores higher on coding benchmarks: 55.5 versus 43.1 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 200K.

How many benchmarks do Grok Build 0.1 and o3-pro share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper