OpenAI, proprietary

o1-mini

o1-mini by OpenAI ranks 235th of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.0. Its strongest category is agentic & tool use, where it ranks 118th.

Last verified

Specifications

Noometry rank
#235 of 354
Index score
34.0
Evidence
Confirmed 39 results
Provider
OpenAI
Released
September 12, 2024
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

o1-mini category scores
  1. Coding 35.5
  2. Agentic & Tool Use 24.6
  3. Reasoning 8.8
  4. Math 35.4
  5. Knowledge 34.9
  6. Multilingual 43.6
  7. Instruction Following 66.7
  8. Long Context 40.1
  9. Writing & Preference 48.4
o1-mini category ranks
CategoryScoreRankResults
Coding35.5#2244
Agentic & Tool Use24.6#1181
Reasoning8.8#3466
Math35.4#1864
Knowledge34.9#1923
Multilingual43.6#1821
Instruction Following66.7#2062
Long Context40.1#1611
Writing & Preference48.4#2025

Strengths and weaknesses

Categories where o1-mini places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

o1-mini: strongest categories
CategoryScorevs medianRank
Long Context40.1−0.8#161 of 296, top 55%
Math35.4−1.2#186 of 327, top 57%
Knowledge34.9−2.5#192 of 314, top 62%

Weakest categories

o1-mini: weakest categories
CategoryScorevs medianRank
Reasoning8.8−14.8#346 of 350, top 99%
Agentic & Tool Use24.6−5.8#118 of 154, top 77%
Instruction Following66.7−4.5#206 of 305, top 68%

Closest competitors

The models ranked just above and below o1-mini. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to o1-mini
ModelRankScoreBlended $/MSpeed
Claude 3.5 Sonnet#23134.6——Compare
Qwen3 Coder Next#23234.3$0.29—Compare
Devstral Small 2505#23334.3$0.1588Compare
Qwen1.5-110B#23434.2——Compare
Qwen3.5-9B#23633.8$0.11—Compare
Codellama 70b Instruct#23733.7——Compare
Qwen3 8B#23833.7$0.31—Compare
Grok-2 (Dec 2024)#23933.7——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

o1-mini Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot32.9%#31 of 44, top 71%Epoch AI
WeirdML36.3%#89 of 119, top 75%mediumEpoch AI
LiveBench Coding48%#18 of 39, top 47%mediumEpoch AI
LMArena Coding1362#164 of 294, top 56%LMArena2026-10-08
HumanEval+89%#2 of 45, top 5%sept 2024EvalPlus
MBPP+78.8%#2 of 38, top 6%sept 2024EvalPlus

Agentic & Tool Use

o1-mini Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench10%#17 of 21, top 81%mediumEpoch AI

Reasoning

o1-mini Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20.8%#70 of 83, top 85%Epoch AI
SimpleBench18.1%#74 of 77, top 97%mediumEpoch AI
ARC-AGI-114%#72 of 83, top 87%Epoch AI
ARC-AGI-114%mediumEpoch AI
LiveBench Reasoning72.3%#9 of 39, top 24%mediumEpoch AI
LMArena Hard Prompts1333#173 of 297, top 59%LMArena2026-10-08
LiveBench Data Analysis57.9%#15 of 39, top 39%mediumEpoch AI
Epoch Capabilities Index135.82#123 of 213, top 58%Epoch AI2024-09-12
LiveBench57.8%#14 of 39, top 36%mediumEpoch AI

Math

o1-mini Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202546.9%#114 of 173, top 66%highEpoch AI2025-03-06
OTIS Mock AIME 2024-202544.7%mediumEpoch AI2025-03-06
LiveBench Math62%#12 of 39, top 31%mediumEpoch AI
LMArena Math1358#161 of 285, top 57%LMArena2026-10-08
MATH Level 589.2%#16 of 79, top 21%highEpoch AI2025-02-13
MATH Level 584.3%mediumEpoch AI2025-01-27
FrontierMath (Feb 2025 set)1.4%highEpoch AI2025-03-06
FrontierMath (Feb 2025 set)1.7%#57 of 68, top 84%mediumEpoch AI2025-03-06

Knowledge

o1-mini Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond62.4%#112 of 186, top 61%highEpoch AI2025-02-13
GPQA Diamond59.5%mediumEpoch AI2025-01-27
Confabulations (lower is better)18.6%#30 of 51, top 59%Lech Mazur benchmarks
LMArena Expert1316#171 of 273, top 63%LMArena2026-10-08

Multilingual

o1-mini Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1289#182 of 297, top 62%LMArena2026-10-08
LMArena Chinese1314#184 of 285, top 65%LMArena2026-10-08
LMArena French1293#164 of 223, top 74%LMArena2026-10-08
LMArena German1278#161 of 231, top 70%LMArena2026-10-08
LMArena Japanese1245#141 of 211, top 67%LMArena2026-10-08
LMArena Korean1223#153 of 213, top 72%LMArena2026-10-08
LMArena Russian1283#190 of 283, top 68%LMArena2026-10-08
LMArena Spanish1303#160 of 226, top 71%LMArena2026-10-08

Instruction Following

o1-mini Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following65.4%#23 of 39, top 59%mediumEpoch AI
LMArena Instruction Following1304#175 of 298, top 59%LMArena2026-10-08

Long Context

o1-mini Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1320#172 of 291, top 60%LMArena2026-10-08

Writing & Preference

o1-mini Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1317#182 of 297, top 62%LMArena2026-10-08
LMArena Creative Writing1244#211 of 295, top 72%LMArena2026-10-08
Short-Story Creative Writing64.9%#34 of 39, top 88%mediumEpoch AI
LMArena Multi-Turn1314#179 of 295, top 61%LMArena2026-10-08
LiveBench Language40.9%#18 of 39, top 47%mediumEpoch AI

Compare o1-mini

Other OpenAI models

Frequently asked questions

How good is o1-mini?

o1-mini by OpenAI ranks 235th of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.0. Its strongest category is agentic & tool use, where it ranks 118th.

Is o1-mini open source?

No. o1-mini is proprietary and available only through OpenAI's API and partner platforms.

What are o1-mini's strengths and weaknesses?

Relative to other ranked models, o1-mini places best in long context, math, knowledge and lowest in reasoning, agentic & tool use, instruction following.

What is o1-mini best at?

Its best category is agentic & tool use, where it ranks 118th on Noometry.