OpenAI, proprietary

o1

o1 by OpenAI ranks 143rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.9. Its strongest category is long context, where it ranks 9th. API pricing starts at $15 per million input tokens and $60 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#143 of 354
Index score
40.9
Evidence
Confirmed 52 results
Provider
OpenAI
Released
September 12, 2024
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
100K
Input price
$15 / M
Output price
$60 / M
Blended price
$26.25 / M
Output speed
Not measured
Value
#210 of 219
Knowledge cutoff
September 2023
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

o1 category scores
  1. Coding 46.1
  2. Agentic & Tool Use 24.6
  3. Reasoning 27.9
  4. Math 36.1
  5. Knowledge 41.5
  6. Multimodal 34.2
  7. Multilingual 48.6
  8. Instruction Following 74.8
  9. Long Context 50.3
  10. Writing & Preference 55.6
o1 category ranks
CategoryScoreRankResults
Coding46.1#705
Agentic & Tool Use24.6#1171
Reasoning27.9#1119
Math36.1#1755
Knowledge41.5#1105
Multimodal34.2#933
Multilingual48.6#1421
Instruction Following74.8#862
Long Context50.3#92
Writing & Preference55.6#1445

Strengths and weaknesses

Categories where o1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

o1: strongest categories
CategoryScorevs medianRank
Long Context50.3+9.4#9 of 296, top 4%
Coding46.1+7.4#70 of 340, top 21%
Instruction Following74.8+3.5#86 of 305, top 29%

Weakest categories

o1: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use24.6−5.8#117 of 154, top 76%
Multimodal34.2−4.3#93 of 128, top 73%
Math36.1−0.5#175 of 327, top 54%

Closest competitors

The models ranked just above and below o1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to o1
ModelRankScoreBlended $/MSpeed
Hunyuan Turbos 20250226#13941.3——Compare
Kimi K2 (Jul 2025)#14041.2$1201Compare
Grok-3 mini#14141.2—10Compare
Claude Opus 4.1#14241.0$30—Compare
Gemini 3.1 Flash Lite#14440.8$0.5610Compare
Claude Sonnet 4#14540.8$631Compare
Qwen2.5-Max#14640.7——Compare
Nemotron 3 Nano 30B A3B#14740.6$0.0875—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

o1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot61.7%#11 of 44, top 25%highEpoch AI
WeirdML47.6%#54 of 119, top 46%Epoch AI
WeirdML46.1%highEpoch AI
LiveBench Coding69.7%#8 of 39, top 21%highEpoch AI
LMArena Coding1367LMArena2026-10-08
LMArena Coding1367#162 of 294, top 56%LMArena2026-10-08
CadEval56%#4 of 14, top 29%mediumEpoch AI
HumanEval+89%Best of 45sept 2024EvalPlus
MBPP+80.2%Best of 38sept 2024EvalPlus

Agentic & Tool Use

o1 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench10%#16 of 21, top 77%Epoch AI
METR Time Horizons45.1%Epoch AI
METR Time Horizons51.1%#23 of 32, top 72%mediumEpoch AI

Reasoning

o1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench41.7%#50 of 77, top 65%Epoch AI
SimpleBench40.1%highEpoch AI
SimpleBench36.7%mediumEpoch AI
ARC-AGI-118%Epoch AI
ARC-AGI-127.2%lowEpoch AI
ARC-AGI-130.7%#65 of 83, top 79%mediumEpoch AI
Chess Puzzles15%#68 of 129, top 53%highEpoch AI2026-08-07
Chess Puzzles7%lowEpoch AI2026-07-15
Chess Puzzles12%mediumEpoch AI2026-08-07
EnigmaEval5.7%#21 of 38, top 56%Epoch AI
LiveBench Reasoning91.6%#2 of 39, top 6%highEpoch AI
LMArena Hard Prompts1371#147 of 297, top 50%LMArena2026-10-08
LMArena Hard Prompts1354LMArena2026-10-08
DTBench74.7%#88 of 151, top 59%highEpoch AI
LiveBench Data Analysis65.5%#9 of 39, top 24%highEpoch AI
LMCA22.3%#90 of 125, top 72%highEpoch AI
Epoch Capabilities Index134.79Epoch AI2024-09-12
Epoch Capabilities Index141.91#100 of 213, top 47%Epoch AI2024-12-17
LiveBench75.7%#5 of 39, top 13%highEpoch AI

Math

o1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)14.7%#74 of 81, top 92%highEpoch AI2026-08-28
FrontierMath (Tiers 1-3)8.4%lowEpoch AI2026-08-27
FrontierMath (Tiers 1-3)10.2%mediumEpoch AI2026-08-27
OTIS Mock AIME 2024-202531.1%Epoch AI2025-03-07
OTIS Mock AIME 2024-202573.3%#86 of 173, top 50%highEpoch AI2026-07-15
OTIS Mock AIME 2024-202553.3%lowEpoch AI2026-07-15
OTIS Mock AIME 2024-202573.3%mediumEpoch AI2025-02-27
LiveBench Math80.3%#4 of 39, top 11%highEpoch AI
LMArena Math1373LMArena2026-10-08
LMArena Math1388#140 of 285, top 50%LMArena2026-10-08
MATH Level 581.6%Epoch AI2025-01-27
MATH Level 594.7%#12 of 79, top 16%highEpoch AI2025-02-13
MATH Level 594.4%mediumEpoch AI2025-01-27
FrontierMath (Feb 2025 set)9.3%#38 of 68, top 56%highEpoch AI2025-03-07

Knowledge

o1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond50.3%Epoch AI2025-01-27
GPQA Diamond76.8%#87 of 186, top 47%highEpoch AI2025-02-13
GPQA Diamond74.2%lowEpoch AI2026-07-20
GPQA Diamond75.8%mediumEpoch AI2025-01-27
Humanity's Last Exam8%#30 of 41, top 74%Epoch AI
SimpleQA Verified41.1%#40 of 77, top 52%highEpoch AI2026-08-31
Confabulations (lower is better)13%Lech Mazur benchmarks
Confabulations (lower is better)11.7%#5 of 51, top 10%medium reasoningLech Mazur benchmarks
LMArena Expert1338LMArena2026-10-08
LMArena Expert1361#147 of 273, top 54%LMArena2026-10-08

Multimodal

o1 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1168#92 of 122, top 76%LMArena2026-10-09
GeoBench80%#5 of 25, top 20%mediumEpoch AI
VPCT37%#18 of 24, top 75%mediumEpoch AI
SpatialViz-Bench41.4%#2 of 8, top 25%Epoch AI

Multilingual

o1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1358#142 of 297, top 48%LMArena2026-10-08
LMArena Non-English1315LMArena2026-10-08
LMArena Chinese1326LMArena2026-10-08
LMArena Chinese1394#140 of 285, top 50%LMArena2026-10-08
LMArena French1344LMArena2026-10-08
LMArena French1344#146 of 223, top 66%LMArena2026-10-08
LMArena German1337#139 of 231, top 61%LMArena2026-10-08
LMArena German1309LMArena2026-10-08
LMArena Japanese1294LMArena2026-10-08
LMArena Japanese1346#100 of 211, top 48%LMArena2026-10-08
LMArena Korean1396#59 of 213, top 28%LMArena2026-10-08
LMArena Korean1290LMArena2026-10-08
LMArena Russian1356#143 of 283, top 51%LMArena2026-10-08
LMArena Russian1316LMArena2026-10-08
LMArena Spanish1302LMArena2026-10-08
LMArena Spanish1345#150 of 226, top 67%LMArena2026-10-08

Instruction Following

o1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following81.5%#7 of 39, top 18%highEpoch AI
LMArena Instruction Following1367#132 of 298, top 45%LMArena2026-10-08
LMArena Instruction Following1342LMArena2026-10-08

Long Context

o1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench83.3%#10 of 47, top 22%mediumEpoch AI
LMArena Longer Query1344LMArena2026-10-08
LMArena Longer Query1378#135 of 291, top 47%LMArena2026-10-08

Writing & Preference

o1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1366#146 of 297, top 50%LMArena2026-10-08
LMArena Text1353LMArena2026-10-08
LMArena Creative Writing1318LMArena2026-10-08
LMArena Creative Writing1348#130 of 295, top 45%LMArena2026-10-08
Short-Story Creative Writing70.2%#31 of 39, top 80%mediumEpoch AI
LMArena Multi-Turn1355LMArena2026-10-08
LMArena Multi-Turn1369#141 of 295, top 48%LMArena2026-10-08
LiveBench Language65.4%#3 of 39, top 8%highEpoch AI

API pricing by provider

o1 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$15$60$7.502026-10-10
openai$15$60$7.502026-10-10
openrouter$15$60$7.502026-10-10

Compare o1

Other OpenAI models

Frequently asked questions

How good is o1?

o1 by OpenAI ranks 143rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.9. Its strongest category is long context, where it ranks 9th. API pricing starts at $15 per million input tokens and $60 per million output tokens, with a 200K-token context window.

How much does o1 cost?

o1 costs $15 per million input tokens and $60 per million output tokens on OpenAI's own API, with cached input at $7.50.

What is o1's context window?

o1 accepts up to 200K tokens of input and can write up to 100K tokens in one response.

Is o1 open source?

No. o1 is proprietary and available only through OpenAI's API and partner platforms.

What are o1's strengths and weaknesses?

Relative to other ranked models, o1 places best in long context, coding, instruction following and lowest in agentic & tool use, multimodal, math.

What is o1 best at?

Its best category is long context, where it ranks 9th on Noometry.