OpenAI, proprietary

o3

o3 by OpenAI ranks 61st of 354 ranked models on the Noometry Index as of October 2026, with a score of 47.5. Its strongest category is long context, where it ranks 6th. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#61 of 354
Index score
47.5
Evidence
Confirmed 63 results
Provider
OpenAI
Released
April 16, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
100K
Input price
$2 / M
Output price
$8 / M
Blended price
$3.50 / M
Output speed
3 tokens/s Kagi
Value
#168 of 219
Knowledge cutoff
May 2024
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

o3 category scores
  1. Coding 46.8
  2. Agentic & Tool Use 34.5
  3. Reasoning 32.0
  4. Math 50.2
  5. Knowledge 54.6
  6. Multimodal 41.4
  7. Multilingual 51.7
  8. Instruction Following 72.8
  9. Long Context 53.3
  10. Writing & Preference 63.5
o3 category ranks
CategoryScoreRankResults
Coding46.8#647
Agentic & Tool Use34.5#444
Reasoning32.0#7811
Math50.2#585
Knowledge54.6#527
Multimodal41.4#363
Multilingual51.7#1051
Instruction Following72.8#1272
Long Context53.3#63
Writing & Preference63.5#646

Strengths and weaknesses

Categories where o3 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

o3: strongest categories
CategoryScorevs medianRank
Long Context53.3+12.3#6 of 296, top 3%
Knowledge54.6+17.3#52 of 314, top 17%
Math50.2+13.6#58 of 327, top 18%

Weakest categories

o3: weakest categories
CategoryScorevs medianRank
Instruction Following72.8+1.5#127 of 305, top 42%
Multilingual51.7+4.3#105 of 297, top 36%
Agentic & Tool Use34.5+4.1#44 of 154, top 29%

Closest competitors

The models ranked just above and below o3. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to o3
ModelRankScoreBlended $/MSpeed
Kimi K2.5#5748.1$0.9066Compare
Step 5 Preview#5847.9$1.43—Compare
GLM-5.1#5947.8$2.15—Compare
Kimi K2.6#6047.7$1.71—Compare
Qwen3.6 Plus#6247.5$1.13—Compare
Inkling-Small#6346.5$0.64—Compare
GPT-5 Pro#6446.4$41.255Compare
Grok 4.20 Multi-Agent#6546.2$1.56—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

o3 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified62.3%#27 of 32, top 85%mediumEpoch AI2026-02-12
SWE-bench Verified (bash only)58.4%#21 of 39, top 54%SWE-bench2025-07-26
Aider Polyglot76.9%Epoch AI
Aider Polyglot81.3%#4 of 44, top 10%highEpoch AI
Aider Polyglot76.9%mediumEpoch AI
GSO8.8%#18 of 31, top 59%highEpoch AI
WeirdML52.4%#44 of 119, top 37%highEpoch AI
LMArena Coding1408#132 of 294, top 45%LMArena2026-10-08
CadEval74%Best of 14mediumEpoch AI
ALE-Bench933.55#46 of 105, top 44%highEpoch AI

Agentic & Tool Use

o3 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard63%#7 of 49, top 15%promptBerkeley Function Calling Leaderboard
GDPval30.8%#7 of 11, top 64%mediumEpoch AI
DeepResearch Bench45.2%#16 of 24, top 67%mediumEpoch AI
OSWorld23%#7 of 8, top 88%mediumEpoch AI
LMArena Search1144#25 of 32, top 79%LMArena2026-08-24
METR Time Horizons63.6%Epoch AI
METR Time Horizons65.4%#14 of 32, top 44%mediumEpoch AI

Reasoning

o3 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-26.5%#51 of 83, top 62%highEpoch AI
ARC-AGI-22%lowEpoch AI
ARC-AGI-23%mediumEpoch AI
SimpleBench53.1%#38 of 77, top 50%highEpoch AI
Kagi LLM Benchmark67.6%#27 of 99, top 28%Kagi LLM Benchmark
ARC-AGI-160.8%#50 of 83, top 61%highEpoch AI
ARC-AGI-141.5%lowEpoch AI
ARC-AGI-153.8%mediumEpoch AI
CritPt1.4%#77 of 134, top 58%highEpoch AI
Chess Puzzles34%highEpoch AI2026-08-07
Chess Puzzles27%lowEpoch AI2026-07-15
Chess Puzzles38%#26 of 129, top 21%mediumEpoch AI2026-08-07
EnigmaEval11.9%highEpoch AI
EnigmaEval13.1%#10 of 38, top 27%mediumEpoch AI
LMArena Hard Prompts1402#124 of 297, top 42%LMArena2026-10-08
Mystery Game Puzzles29%#29 of 74, top 40%highEpoch AI2026-08-27
Mystery Game Puzzles23%mediumEpoch AI2026-08-27
DTBench84.8%#54 of 151, top 36%highEpoch AI
LMCA39.7%#46 of 125, top 37%highEpoch AI
Epoch Capabilities Index146.86#67 of 213, top 32%Epoch AI2025-04-16
ForecastBench62.5Best of 72Epoch AI

Math

o3 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)33.3%#60 of 81, top 75%highEpoch AI2026-08-27
FrontierMath (Tiers 1-3)19.3%lowEpoch AI2026-08-27
FrontierMath (Tiers 1-3)29.8%mediumEpoch AI2026-08-27
OTIS Mock AIME 2024-202583.9%highEpoch AI2025-04-16
OTIS Mock AIME 2024-202560%lowEpoch AI2026-07-15
OTIS Mock AIME 2024-202584.4%#71 of 173, top 42%mediumEpoch AI2026-08-07
Omni-MATH71.4%#4 of 57, top 8%HELM Capabilities
LMArena Math1426#93 of 285, top 33%LMArena2026-10-08
MATH Level 597.8%#4 of 79, top 6%highEpoch AI2025-04-16
FrontierMath (Feb 2025 set)18.7%#32 of 68, top 48%highEpoch AI2025-11-16
FrontierMath (Feb 2025 set)9.7%lowEpoch AI2025-11-17
FrontierMath (Feb 2025 set)16.9%mediumEpoch AI2025-11-16
FrontierMath Tier 4 (v1)2.1%#42 of 55, top 77%highEpoch AI2025-07-01

Knowledge

o3 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond81.8%#78 of 186, top 42%highEpoch AI2025-04-16
GPQA Diamond79.8%lowEpoch AI2026-07-15
GPQA Diamond80.8%mediumEpoch AI2026-08-07
Humanity's Last Exam20.3%#18 of 41, top 44%highEpoch AI
Humanity's Last Exam19.2%mediumEpoch AI
SimpleQA Verified49.4%#26 of 77, top 34%highEpoch AI2026-08-27
MMLU-Pro85.9%#6 of 58, top 11%HELM Capabilities
Confabulations (lower is better)14.4%#16 of 51, top 32%high reasoningLech Mazur benchmarks
GPQA (HELM)75.3%#4 of 57, top 8%HELM Capabilities
LMArena Expert1402#120 of 273, top 44%LMArena2026-10-08

Multimodal

o3 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1214#73 of 122, top 60%LMArena2026-10-09
GeoBench60%highEpoch AI
GeoBench74%#10 of 25, top 40%mediumEpoch AI
VPCT52%#7 of 24, top 30%mediumEpoch AI

Multilingual

o3 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1401#105 of 297, top 36%LMArena2026-10-08
LMArena Chinese1437#114 of 285, top 40%LMArena2026-10-08
LMArena French1430#92 of 223, top 42%LMArena2026-10-08
LMArena German1420#77 of 231, top 34%LMArena2026-10-08
LMArena Japanese1403#59 of 211, top 28%LMArena2026-10-08
LMArena Korean1370#87 of 213, top 41%LMArena2026-10-08
LMArena Russian1406#98 of 283, top 35%LMArena2026-10-08
LMArena Spanish1395#116 of 226, top 52%LMArena2026-10-08

Instruction Following

o3 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval86.9%#16 of 57, top 29%HELM Capabilities
LMArena Instruction Following1368#131 of 298, top 44%LMArena2026-10-08

Long Context

o3 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench88.9%#6 of 47, top 13%mediumEpoch AI
CL-bench17.8%#12 of 19, top 64%highEpoch AI
LMArena Longer Query1372#139 of 291, top 48%LMArena2026-10-08

Writing & Preference

o3 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1410#110 of 297, top 38%LMArena2026-10-08
LMArena Creative Writing1359#122 of 295, top 42%LMArena2026-10-08
Short-Story Creative Writing83.9%#5 of 39, top 13%mediumEpoch AI
EQ-Bench Creative Writing1676#35 of 115, top 31%EQ-Bench
WildBench86.1%#4 of 57, top 8%HELM Capabilities
LMArena Multi-Turn1405#117 of 295, top 40%LMArena2026-10-08

API pricing by provider

o3 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$2$8$0.502026-10-10
openai$2$8$0.502026-10-10
openrouter$2$8$0.502026-10-10

Compare o3

Other OpenAI models

Frequently asked questions

How good is o3?

o3 by OpenAI ranks 61st of 354 ranked models on the Noometry Index as of October 2026, with a score of 47.5. Its strongest category is long context, where it ranks 6th. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 200K-token context window.

How much does o3 cost?

o3 costs $2 per million input tokens and $8 per million output tokens on OpenAI's own API, with cached input at $0.50.

What is o3's context window?

o3 accepts up to 200K tokens of input and can write up to 100K tokens in one response.

Is o3 open source?

No. o3 is proprietary and available only through OpenAI's API and partner platforms.

How fast is o3?

o3 generated about 3 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are o3's strengths and weaknesses?

Relative to other ranked models, o3 places best in long context, knowledge, math and lowest in instruction following, multilingual, agentic & tool use.

What is o3 best at?

Its best category is long context, where it ranks 6th on Noometry.