OpenAI, proprietary
o1-mini
o1-mini by OpenAI ranks 235th of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.0. Its strongest category is agentic & tool use, where it ranks 118th.
Last verified
Specifications
- Noometry rank
- #235 of 354
- Index score
- 34.0
- Evidence
- Confirmed 39 results
- Provider
- OpenAI
- Released
- September 12, 2024
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 35.5
- Agentic & Tool Use 24.6
- Reasoning 8.8
- Math 35.4
- Knowledge 34.9
- Multilingual 43.6
- Instruction Following 66.7
- Long Context 40.1
- Writing & Preference 48.4
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 35.5 | #224 | 4 |
| Agentic & Tool Use | 24.6 | #118 | 1 |
| Reasoning | 8.8 | #346 | 6 |
| Math | 35.4 | #186 | 4 |
| Knowledge | 34.9 | #192 | 3 |
| Multilingual | 43.6 | #182 | 1 |
| Instruction Following | 66.7 | #206 | 2 |
| Long Context | 40.1 | #161 | 1 |
| Writing & Preference | 48.4 | #202 | 5 |
Strengths and weaknesses
Categories where o1-mini places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 40.1 | −0.8 | #161 of 296, top 55% |
| Math | 35.4 | −1.2 | #186 of 327, top 57% |
| Knowledge | 34.9 | −2.5 | #192 of 314, top 62% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 8.8 | −14.8 | #346 of 350, top 99% |
| Agentic & Tool Use | 24.6 | −5.8 | #118 of 154, top 77% |
| Instruction Following | 66.7 | −4.5 | #206 of 305, top 68% |
Closest competitors
The models ranked just above and below o1-mini. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude 3.5 Sonnet | #231 | 34.6 | — | — | Compare |
| Qwen3 Coder Next | #232 | 34.3 | $0.29 | — | Compare |
| Devstral Small 2505 | #233 | 34.3 | $0.15 | 88 | Compare |
| Qwen1.5-110B | #234 | 34.2 | — | — | Compare |
| Qwen3.5-9B | #236 | 33.8 | $0.11 | — | Compare |
| Codellama 70b Instruct | #237 | 33.7 | — | — | Compare |
| Qwen3 8B | #238 | 33.7 | $0.31 | — | Compare |
| Grok-2 (Dec 2024) | #239 | 33.7 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 32.9% | #31 of 44, top 71% | Epoch AI | ||
| WeirdML | 36.3% | #89 of 119, top 75% | medium | Epoch AI | |
| LiveBench Coding | 48% | #18 of 39, top 47% | medium | Epoch AI | |
| LMArena Coding | 1362 | #164 of 294, top 56% | LMArena | 2026-10-08 | |
| HumanEval+ | 89% | #2 of 45, top 5% | sept 2024 | EvalPlus | |
| MBPP+ | 78.8% | #2 of 38, top 6% | sept 2024 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Cybench | 10% | #17 of 21, top 81% | medium | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0.8% | #70 of 83, top 85% | Epoch AI | ||
| SimpleBench | 18.1% | #74 of 77, top 97% | medium | Epoch AI | |
| ARC-AGI-1 | 14% | #72 of 83, top 87% | Epoch AI | ||
| ARC-AGI-1 | 14% | medium | Epoch AI | ||
| LiveBench Reasoning | 72.3% | #9 of 39, top 24% | medium | Epoch AI | |
| LMArena Hard Prompts | 1333 | #173 of 297, top 59% | LMArena | 2026-10-08 | |
| LiveBench Data Analysis | 57.9% | #15 of 39, top 39% | medium | Epoch AI | |
| Epoch Capabilities Index | 135.82 | #123 of 213, top 58% | Epoch AI | 2024-09-12 | |
| LiveBench | 57.8% | #14 of 39, top 36% | medium | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 46.9% | #114 of 173, top 66% | high | Epoch AI | 2025-03-06 |
| OTIS Mock AIME 2024-2025 | 44.7% | medium | Epoch AI | 2025-03-06 | |
| LiveBench Math | 62% | #12 of 39, top 31% | medium | Epoch AI | |
| LMArena Math | 1358 | #161 of 285, top 57% | LMArena | 2026-10-08 | |
| MATH Level 5 | 89.2% | #16 of 79, top 21% | high | Epoch AI | 2025-02-13 |
| MATH Level 5 | 84.3% | medium | Epoch AI | 2025-01-27 | |
| FrontierMath (Feb 2025 set) | 1.4% | high | Epoch AI | 2025-03-06 | |
| FrontierMath (Feb 2025 set) | 1.7% | #57 of 68, top 84% | medium | Epoch AI | 2025-03-06 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 62.4% | #112 of 186, top 61% | high | Epoch AI | 2025-02-13 |
| GPQA Diamond | 59.5% | medium | Epoch AI | 2025-01-27 | |
| Confabulations (lower is better) | 18.6% | #30 of 51, top 59% | Lech Mazur benchmarks | ||
| LMArena Expert | 1316 | #171 of 273, top 63% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1289 | #182 of 297, top 62% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1314 | #184 of 285, top 65% | LMArena | 2026-10-08 | |
| LMArena French | 1293 | #164 of 223, top 74% | LMArena | 2026-10-08 | |
| LMArena German | 1278 | #161 of 231, top 70% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1245 | #141 of 211, top 67% | LMArena | 2026-10-08 | |
| LMArena Korean | 1223 | #153 of 213, top 72% | LMArena | 2026-10-08 | |
| LMArena Russian | 1283 | #190 of 283, top 68% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1303 | #160 of 226, top 71% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 65.4% | #23 of 39, top 59% | medium | Epoch AI | |
| LMArena Instruction Following | 1304 | #175 of 298, top 59% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1320 | #172 of 291, top 60% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1317 | #182 of 297, top 62% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1244 | #211 of 295, top 72% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 64.9% | #34 of 39, top 88% | medium | Epoch AI | |
| LMArena Multi-Turn | 1314 | #179 of 295, top 61% | LMArena | 2026-10-08 | |
| LiveBench Language | 40.9% | #18 of 39, top 47% | medium | Epoch AI |
Compare o1-mini
- o1-mini vs Qwen1.5-110B
- o1-mini vs Qwen3.5-9B
- o1-mini vs Devstral Small 2505
- o1-mini vs Codellama 70b Instruct
- o1-mini vs Qwen3 Coder Next
- o1-mini vs Qwen3 8B
- o1-mini vs Claude Fable 5.1
- o1-mini vs Gemini 3.8 Flash
- o1-mini vs Kimi K3
- o1-mini vs Grok 4.6
- o1-mini vs Qwen3.8 Max
- o1-mini vs GLM-5.3
- o1-mini vs Muse Spark 1.3
- o1-mini vs DeepSeek V4 Pro
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is o1-mini?
o1-mini by OpenAI ranks 235th of 354 ranked models on the Noometry Index as of October 2026, with a score of 34.0. Its strongest category is agentic & tool use, where it ranks 118th.
Is o1-mini open source?
No. o1-mini is proprietary and available only through OpenAI's API and partner platforms.
What are o1-mini's strengths and weaknesses?
Relative to other ranked models, o1-mini places best in long context, math, knowledge and lowest in reasoning, agentic & tool use, instruction following.
What is o1-mini best at?
Its best category is agentic & tool use, where it ranks 118th on Noometry.