OpenAI, proprietary
o1
o1 by OpenAI ranks 143rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.9. Its strongest category is long context, where it ranks 9th. API pricing starts at $15 per million input tokens and $60 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #143 of 354
- Index score
- 40.9
- Evidence
- Confirmed 52 results
- Provider
- OpenAI
- Released
- September 12, 2024
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 100K
- Input price
- $15 / M
- Output price
- $60 / M
- Blended price
- $26.25 / M
- Output speed
- Not measured
- Value
- #210 of 219
- Knowledge cutoff
- September 2023
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 46.1
- Agentic & Tool Use 24.6
- Reasoning 27.9
- Math 36.1
- Knowledge 41.5
- Multimodal 34.2
- Multilingual 48.6
- Instruction Following 74.8
- Long Context 50.3
- Writing & Preference 55.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 46.1 | #70 | 5 |
| Agentic & Tool Use | 24.6 | #117 | 1 |
| Reasoning | 27.9 | #111 | 9 |
| Math | 36.1 | #175 | 5 |
| Knowledge | 41.5 | #110 | 5 |
| Multimodal | 34.2 | #93 | 3 |
| Multilingual | 48.6 | #142 | 1 |
| Instruction Following | 74.8 | #86 | 2 |
| Long Context | 50.3 | #9 | 2 |
| Writing & Preference | 55.6 | #144 | 5 |
Strengths and weaknesses
Categories where o1 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 50.3 | +9.4 | #9 of 296, top 4% |
| Coding | 46.1 | +7.4 | #70 of 340, top 21% |
| Instruction Following | 74.8 | +3.5 | #86 of 305, top 29% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 24.6 | −5.8 | #117 of 154, top 76% |
| Multimodal | 34.2 | −4.3 | #93 of 128, top 73% |
| Math | 36.1 | −0.5 | #175 of 327, top 54% |
Closest competitors
The models ranked just above and below o1. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Hunyuan Turbos 20250226 | #139 | 41.3 | — | — | Compare |
| Kimi K2 (Jul 2025) | #140 | 41.2 | $1 | 201 | Compare |
| Grok-3 mini | #141 | 41.2 | — | 10 | Compare |
| Claude Opus 4.1 | #142 | 41.0 | $30 | — | Compare |
| Gemini 3.1 Flash Lite | #144 | 40.8 | $0.56 | 10 | Compare |
| Claude Sonnet 4 | #145 | 40.8 | $6 | 31 | Compare |
| Qwen2.5-Max | #146 | 40.7 | — | — | Compare |
| Nemotron 3 Nano 30B A3B | #147 | 40.6 | $0.0875 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 61.7% | #11 of 44, top 25% | high | Epoch AI | |
| WeirdML | 47.6% | #54 of 119, top 46% | Epoch AI | ||
| WeirdML | 46.1% | high | Epoch AI | ||
| LiveBench Coding | 69.7% | #8 of 39, top 21% | high | Epoch AI | |
| LMArena Coding | 1367 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1367 | #162 of 294, top 56% | LMArena | 2026-10-08 | |
| CadEval | 56% | #4 of 14, top 29% | medium | Epoch AI | |
| HumanEval+ | 89% | Best of 45 | sept 2024 | EvalPlus | |
| MBPP+ | 80.2% | Best of 38 | sept 2024 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Cybench | 10% | #16 of 21, top 77% | Epoch AI | ||
| METR Time Horizons | 45.1% | Epoch AI | |||
| METR Time Horizons | 51.1% | #23 of 32, top 72% | medium | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 41.7% | #50 of 77, top 65% | Epoch AI | ||
| SimpleBench | 40.1% | high | Epoch AI | ||
| SimpleBench | 36.7% | medium | Epoch AI | ||
| ARC-AGI-1 | 18% | Epoch AI | |||
| ARC-AGI-1 | 27.2% | low | Epoch AI | ||
| ARC-AGI-1 | 30.7% | #65 of 83, top 79% | medium | Epoch AI | |
| Chess Puzzles | 15% | #68 of 129, top 53% | high | Epoch AI | 2026-08-07 |
| Chess Puzzles | 7% | low | Epoch AI | 2026-07-15 | |
| Chess Puzzles | 12% | medium | Epoch AI | 2026-08-07 | |
| EnigmaEval | 5.7% | #21 of 38, top 56% | Epoch AI | ||
| LiveBench Reasoning | 91.6% | #2 of 39, top 6% | high | Epoch AI | |
| LMArena Hard Prompts | 1371 | #147 of 297, top 50% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1354 | LMArena | 2026-10-08 | ||
| DTBench | 74.7% | #88 of 151, top 59% | high | Epoch AI | |
| LiveBench Data Analysis | 65.5% | #9 of 39, top 24% | high | Epoch AI | |
| LMCA | 22.3% | #90 of 125, top 72% | high | Epoch AI | |
| Epoch Capabilities Index | 134.79 | Epoch AI | 2024-09-12 | ||
| Epoch Capabilities Index | 141.91 | #100 of 213, top 47% | Epoch AI | 2024-12-17 | |
| LiveBench | 75.7% | #5 of 39, top 13% | high | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 14.7% | #74 of 81, top 92% | high | Epoch AI | 2026-08-28 |
| FrontierMath (Tiers 1-3) | 8.4% | low | Epoch AI | 2026-08-27 | |
| FrontierMath (Tiers 1-3) | 10.2% | medium | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 31.1% | Epoch AI | 2025-03-07 | ||
| OTIS Mock AIME 2024-2025 | 73.3% | #86 of 173, top 50% | high | Epoch AI | 2026-07-15 |
| OTIS Mock AIME 2024-2025 | 53.3% | low | Epoch AI | 2026-07-15 | |
| OTIS Mock AIME 2024-2025 | 73.3% | medium | Epoch AI | 2025-02-27 | |
| LiveBench Math | 80.3% | #4 of 39, top 11% | high | Epoch AI | |
| LMArena Math | 1373 | LMArena | 2026-10-08 | ||
| LMArena Math | 1388 | #140 of 285, top 50% | LMArena | 2026-10-08 | |
| MATH Level 5 | 81.6% | Epoch AI | 2025-01-27 | ||
| MATH Level 5 | 94.7% | #12 of 79, top 16% | high | Epoch AI | 2025-02-13 |
| MATH Level 5 | 94.4% | medium | Epoch AI | 2025-01-27 | |
| FrontierMath (Feb 2025 set) | 9.3% | #38 of 68, top 56% | high | Epoch AI | 2025-03-07 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 50.3% | Epoch AI | 2025-01-27 | ||
| GPQA Diamond | 76.8% | #87 of 186, top 47% | high | Epoch AI | 2025-02-13 |
| GPQA Diamond | 74.2% | low | Epoch AI | 2026-07-20 | |
| GPQA Diamond | 75.8% | medium | Epoch AI | 2025-01-27 | |
| Humanity's Last Exam | 8% | #30 of 41, top 74% | Epoch AI | ||
| SimpleQA Verified | 41.1% | #40 of 77, top 52% | high | Epoch AI | 2026-08-31 |
| Confabulations (lower is better) | 13% | Lech Mazur benchmarks | |||
| Confabulations (lower is better) | 11.7% | #5 of 51, top 10% | medium reasoning | Lech Mazur benchmarks | |
| LMArena Expert | 1338 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1361 | #147 of 273, top 54% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1168 | #92 of 122, top 76% | LMArena | 2026-10-09 | |
| GeoBench | 80% | #5 of 25, top 20% | medium | Epoch AI | |
| VPCT | 37% | #18 of 24, top 75% | medium | Epoch AI | |
| SpatialViz-Bench | 41.4% | #2 of 8, top 25% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1358 | #142 of 297, top 48% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1315 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1326 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1394 | #140 of 285, top 50% | LMArena | 2026-10-08 | |
| LMArena French | 1344 | LMArena | 2026-10-08 | ||
| LMArena French | 1344 | #146 of 223, top 66% | LMArena | 2026-10-08 | |
| LMArena German | 1337 | #139 of 231, top 61% | LMArena | 2026-10-08 | |
| LMArena German | 1309 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1294 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1346 | #100 of 211, top 48% | LMArena | 2026-10-08 | |
| LMArena Korean | 1396 | #59 of 213, top 28% | LMArena | 2026-10-08 | |
| LMArena Korean | 1290 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1356 | #143 of 283, top 51% | LMArena | 2026-10-08 | |
| LMArena Russian | 1316 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1302 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1345 | #150 of 226, top 67% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 81.5% | #7 of 39, top 18% | high | Epoch AI | |
| LMArena Instruction Following | 1367 | #132 of 298, top 45% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1342 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 83.3% | #10 of 47, top 22% | medium | Epoch AI | |
| LMArena Longer Query | 1344 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1378 | #135 of 291, top 47% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1366 | #146 of 297, top 50% | LMArena | 2026-10-08 | |
| LMArena Text | 1353 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1318 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1348 | #130 of 295, top 45% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 70.2% | #31 of 39, top 80% | medium | Epoch AI | |
| LMArena Multi-Turn | 1355 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1369 | #141 of 295, top 48% | LMArena | 2026-10-08 | |
| LiveBench Language | 65.4% | #3 of 39, top 8% | high | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $15 | $60 | $7.50 | 2026-10-10 |
| openai | $15 | $60 | $7.50 | 2026-10-10 |
| openrouter | $15 | $60 | $7.50 | 2026-10-10 |
Compare o1
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is o1?
o1 by OpenAI ranks 143rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.9. Its strongest category is long context, where it ranks 9th. API pricing starts at $15 per million input tokens and $60 per million output tokens, with a 200K-token context window.
How much does o1 cost?
o1 costs $15 per million input tokens and $60 per million output tokens on OpenAI's own API, with cached input at $7.50.
What is o1's context window?
o1 accepts up to 200K tokens of input and can write up to 100K tokens in one response.
Is o1 open source?
No. o1 is proprietary and available only through OpenAI's API and partner platforms.
What are o1's strengths and weaknesses?
Relative to other ranked models, o1 places best in long context, coding, instruction following and lowest in agentic & tool use, multimodal, math.
What is o1 best at?
Its best category is long context, where it ranks 9th on Noometry.