Model comparison
Grok Build 0.1 vs o3-pro
o3-pro is the stronger model overall, scoring 42.9 to 36.4 on the Noometry Index. Grok Build 0.1 costs 28× less per token, which makes it the better buy when o3-pro's lead doesn't matter for your workload.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in coding, where o3-pro leads 55.5 to 43.1.
- Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $20 / $80 for o3-pro.
- Grok Build 0.1 accepts more context: 256K tokens versus 200K.
Side by side
| Grok Build 0.1 | o3-pro | |
|---|---|---|
| Provider | xAI | OpenAI |
| Noometry Index | 36.4 | 42.9 |
| Released | 2026-04-16 | 2025-06-10 |
| Weights | Proprietary | Proprietary |
| Context window | 256K | 200K |
| Max output | 256K | 100K |
| Input $ / M tokens | $1 | $20 |
| Output $ / M tokens | $2 | $80 |
| Results tracked | 3 | 12 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding o3-pro leads
Grok Build 0.1: 43.1 (#91), o3-pro: 55.5 (#24)
| Benchmark | Grok Build 0.1 | o3-pro |
|---|---|---|
| Aider Polyglot | — | 84.9% |
| SciCode | 50.2% | — |
| WeirdML | — | 58.2% |
Agentic & Tool Use Not comparable
Grok Build 0.1: 22.7 (#129), o3-pro: —
| Benchmark | Grok Build 0.1 | o3-pro |
|---|---|---|
| GBAEval | 2.4% | — |
Reasoning Grok Build 0.1 leads
Grok Build 0.1: 32.2 (#77), o3-pro: 23.8 (#171)
| Benchmark | Grok Build 0.1 | o3-pro |
|---|---|---|
| ARC-AGI-2 | — | 4.9% |
| Kagi LLM Benchmark | — | 72.1% |
| ARC-AGI-1 | — | 59.3% |
| CritPt | 9.1% | — |
| DTBench | — | 86.9% |
| LMCA | — | 38.5% |
| Epoch Capabilities Index | — | 147.42 |
Knowledge Not comparable
Grok Build 0.1: —, o3-pro: 29.5 (#238)
| Benchmark | Grok Build 0.1 | o3-pro |
|---|---|---|
| Confabulations | — | 14.2% |
| Vectara Hallucination Rate | — | 23.3% |
Long Context Not comparable
Grok Build 0.1: —, o3-pro: 72.2 (#1)
| Benchmark | Grok Build 0.1 | o3-pro |
|---|---|---|
| Fiction.LiveBench | — | 97.2% |
Writing & Preference Not comparable
Grok Build 0.1: —, o3-pro: 57.1 (#133)
| Benchmark | Grok Build 0.1 | o3-pro |
|---|---|---|
| Short-Story Creative Writing | — | 84.4% |
Frequently asked questions
Is Grok Build 0.1 better than o3-pro?
o3-pro is the stronger model overall, scoring 42.9 to 36.4 on the Noometry Index. Grok Build 0.1 costs 28× less per token, which makes it the better buy when o3-pro's lead doesn't matter for your workload.
Which is cheaper, Grok Build 0.1 or o3-pro?
Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; o3-pro lists at $20 and $80.
Is Grok Build 0.1 or o3-pro better for coding?
o3-pro scores higher on coding benchmarks: 55.5 versus 43.1 in the Noometry coding category.
Which has the bigger context window?
Grok Build 0.1 does, with 256K tokens against 200K.
How many benchmarks do Grok Build 0.1 and o3-pro share?
0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and o3-pro has 12.