OpenAI, proprietary
o3-pro
o3-pro by OpenAI ranks 105th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.9. Its strongest category is long context, where it ranks 1st. API pricing starts at $20 per million input tokens and $80 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #105 of 354
- Index score
- 42.9
- Evidence
- Confirmed 12 results
- Provider
- OpenAI
- Released
- June 10, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 100K
- Input price
- $20 / M
- Output price
- $80 / M
- Blended price
- $35 / M
- Output speed
- 1 tokens/s Kagi
- Value
- #213 of 219
- Knowledge cutoff
- May 2024
- Input
- text, image
Category scores
Each category score combines every public result we have in that category.
- Coding 55.5
- Reasoning 23.8
- Knowledge 29.5
- Long Context 72.2
- Writing & Preference 57.1
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 55.5 | #24 | 2 |
| Reasoning | 23.8 | #171 | 5 |
| Knowledge | 29.5 | #238 | 2 |
| Long Context | 72.2 | #1 | 1 |
| Writing & Preference | 57.1 | #133 | 1 |
Strengths and weaknesses
Categories where o3-pro places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 72.2 | +31.2 | #1 of 296, top 1% |
| Coding | 55.5 | +16.8 | #24 of 340, top 8% |
Closest competitors
The models ranked just above and below o3-pro. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Amazon Nova Experimental Chat 11 10 | #101 | 43.0 | — | — | Compare |
| Qwen3-Next 80B-A3B Instruct | #102 | 43.0 | $0.88 | 111 | Compare |
| MiMo-V2-Pro | #103 | 43.0 | $0.54 | — | Compare |
| Amazon Nova Experimental Chat 12 10 | #104 | 42.9 | — | — | Compare |
| Qwen3.5 Plus | #106 | 42.9 | $0.90 | — | Compare |
| Amazon Nova Experimental Chat 26 01 10 | #107 | 42.8 | — | — | Compare |
| DeepSeek-V3.1 | #108 | 42.8 | $0.42 | 328 | Compare |
| GPT-5.3 Chat | #109 | 42.8 | $4.81 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 84.9% | #2 of 44, top 5% | high | Epoch AI | |
| WeirdML | 58.2% | #34 of 119, top 29% | high | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 4.9% | #59 of 83, top 72% | high | Epoch AI | |
| ARC-AGI-2 | 2% | low | Epoch AI | ||
| ARC-AGI-2 | 1.9% | medium | Epoch AI | ||
| Kagi LLM Benchmark | 72.1% | #21 of 99, top 22% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 59.3% | #51 of 83, top 62% | high | Epoch AI | |
| ARC-AGI-1 | 44.3% | low | Epoch AI | ||
| ARC-AGI-1 | 57% | medium | Epoch AI | ||
| DTBench | 86.9% | #53 of 151, top 36% | high | Epoch AI | |
| LMCA | 38.5% | #50 of 125, top 40% | high | Epoch AI | |
| Epoch Capabilities Index | 147.42 | #64 of 213, top 31% | Epoch AI | 2025-06-10 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Confabulations (lower is better) | 14.2% | #15 of 51, top 30% | medium reasoning | Lech Mazur benchmarks | |
| Vectara Hallucination Rate (lower is better) | 23.3% | #95 of 96, top 99% | Vectara Hallucination Leaderboard |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 97.2% | #2 of 47, top 5% | medium | Epoch AI |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Short-Story Creative Writing | 84.4% | #4 of 39, top 11% | medium | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| openai | $20 | $80 | — | 2026-10-10 |
| openrouter | $20 | $80 | — | 2026-10-10 |
Compare o3-pro
- o3-pro vs o1-pro
- o3-pro vs Amazon Nova Experimental Chat 12 10
- o3-pro vs Qwen3.5 Plus
- o3-pro vs MiMo-V2-Pro
- o3-pro vs Amazon Nova Experimental Chat 26 01 10
- o3-pro vs Qwen3-Next 80B-A3B Instruct
- o3-pro vs DeepSeek-V3.1
- o3-pro vs Claude Fable 5.1
- o3-pro vs Gemini 3.8 Flash
- o3-pro vs Kimi K3
- o3-pro vs Grok 4.6
- o3-pro vs Qwen3.8 Max
- o3-pro vs GLM-5.3
- o3-pro vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is o3-pro?
o3-pro by OpenAI ranks 105th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.9. Its strongest category is long context, where it ranks 1st. API pricing starts at $20 per million input tokens and $80 per million output tokens, with a 200K-token context window.
How much does o3-pro cost?
o3-pro costs $20 per million input tokens and $80 per million output tokens on OpenAI's own API.
What is o3-pro's context window?
o3-pro accepts up to 200K tokens of input and can write up to 100K tokens in one response.
Is o3-pro open source?
No. o3-pro is proprietary and available only through OpenAI's API and partner platforms.
How fast is o3-pro?
o3-pro generated about 1 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are o3-pro's strengths and weaknesses?
Relative to other ranked models, o3-pro places best in long context, coding and lowest in knowledge, reasoning.
What is o3-pro best at?
Its best category is long context, where it ranks 1st on Noometry.