OpenAI, proprietary
GPT-4o
GPT-4o by OpenAI ranks 324th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.6. Its strongest category is multimodal, where it ranks 91st. API pricing starts at $2.50 per million input tokens and $10 per million output tokens, with a 128K-token context window.
Last verified
Specifications
- Noometry rank
- #324 of 354
- Index score
- 28.6
- Evidence
- Confirmed 72 results
- Provider
- OpenAI
- Released
- May 13, 2024
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 128K
- Max output
- 16K
- Input price
- $2.50 / M
- Output price
- $10 / M
- Blended price
- $4.38 / M
- Output speed
- Not measured
- Value
- #200 of 219
- Knowledge cutoff
- September 2023
- Input
- text, image
Category scores
Each category score combines every public result we have in that category.
- Coding 24.8
- Agentic & Tool Use 21.0
- Reasoning 9.4
- Math 10.6
- Knowledge 28.8
- Multimodal 34.5
- Multilingual 43.2
- Instruction Following 66.6
- Long Context 39.4
- Writing & Preference 52.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 24.8 | #328 | 10 |
| Agentic & Tool Use | 21.0 | #141 | 4 |
| Reasoning | 9.4 | #343 | 11 |
| Math | 10.6 | #312 | 6 |
| Knowledge | 28.8 | #242 | 8 |
| Multimodal | 34.5 | #91 | 4 |
| Multilingual | 43.2 | #186 | 1 |
| Instruction Following | 66.6 | #207 | 3 |
| Long Context | 39.4 | #179 | 2 |
| Writing & Preference | 52.6 | #166 | 6 |
Strengths and weaknesses
Categories where GPT-4o places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 52.6 | −1.2 | #166 of 312, top 54% |
| Long Context | 39.4 | −1.5 | #179 of 296, top 61% |
| Multilingual | 43.2 | −4.2 | #186 of 297, top 63% |
Closest competitors
The models ranked just above and below GPT-4o. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen2.5 7B Instruct | #320 | 29.0 | $0.31 | — | Compare |
| Llama 3.2 3B | #321 | 28.9 | $0.12 | — | Compare |
| Qwen1.5 4b Chat | #322 | 28.8 | — | — | Compare |
| Llama 3-70B | #323 | 28.8 | — | 104 | Compare |
| Ministral 8B | #325 | 28.2 | $0.15 | — | Compare |
| Gemma 3 4B | #326 | 28.1 | $0.05 | 72 | Compare |
| GPT-4.1 nano | #327 | 27.9 | $0.18 | 135 | Compare |
| Phi 3 Mini 4k Instruct | #328 | 27.9 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 31% | #32 of 32, top 100% | Epoch AI | 2026-02-11 | |
| SWE-bench Verified (bash only) | 21.6% | #35 of 39, top 90% | SWE-bench | 2025-07-20 | |
| Aider Polyglot | 45.3% | #24 of 44, top 55% | Epoch AI | ||
| Aider Polyglot | 27.1% | Epoch AI | |||
| Aider Polyglot | 18.2% | Epoch AI | |||
| Aider Polyglot | 23.1% | Epoch AI | |||
| GSO | 0% | #31 of 31, top 100% | Epoch AI | ||
| WeirdML | 25.1% | #99 of 119, top 84% | Epoch AI | ||
| BigCodeBench Instruct | 51.1% | Best of 64 | BigCodeBench | 2024-05-13 | |
| BigCodeBench Instruct | 48% | BigCodeBench | 2024-11-20 | ||
| LiveBench Coding | 51.4% | #16 of 39, top 42% | Epoch AI | ||
| LiveBench Coding | 46.1% | Epoch AI | |||
| LMArena Coding | 1297 | #199 of 294, top 68% | LMArena | 2026-10-08 | |
| LMArena Coding | 1283 | LMArena | 2026-10-08 | ||
| BigCodeBench Complete | 58.9% | BigCodeBench | 2024-11-20 | ||
| BigCodeBench Complete | 61.1% | #3 of 66, top 5% | BigCodeBench | 2024-05-13 | |
| CadEval | 26% | #12 of 14, top 86% | Epoch AI | ||
| HumanEval+ | 87.2% | #3 of 45, top 7% | aug 2024 | EvalPlus | |
| MBPP+ | 72.2% | #11 of 38, top 29% | aug 2024 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GDPval | 9.9% | #11 of 11, top 100% | Epoch AI | ||
| TheAgentCompany | 8.6% | #8 of 14, top 58% | Epoch AI | ||
| Cybench | 12.5% | #14 of 21, top 67% | Epoch AI | ||
| BALROG | 32.3% | #16 of 35, top 46% | Epoch AI | ||
| LMArena Search | 1006 | #32 of 32, top 100% | LMArena | 2026-08-24 | |
| METR Time Horizons | 40.8% | #26 of 32, top 82% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0% | #77 of 83, top 93% | Epoch AI | ||
| SimpleBench | 17.8% | #75 of 77, top 98% | Epoch AI | ||
| ARC-AGI-1 | 4.5% | #79 of 83, top 96% | Epoch AI | ||
| CritPt | 0% | #112 of 134, top 84% | Epoch AI | ||
| Chess Puzzles | 13% | #73 of 129, top 57% | Epoch AI | 2026-07-15 | |
| EnigmaEval | 0.8% | #35 of 38, top 93% | Epoch AI | ||
| LiveBench Reasoning | 53.9% | Epoch AI | |||
| LiveBench Reasoning | 55.8% | #15 of 39, top 39% | Epoch AI | ||
| LMArena Hard Prompts | 1281 | #199 of 297, top 68% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1264 | LMArena | 2026-10-08 | ||
| DTBench | 64.5% | #103 of 151, top 69% | Epoch AI | ||
| LiveBench Data Analysis | 60.9% | #14 of 39, top 36% | Epoch AI | ||
| LiveBench Data Analysis | 56.1% | Epoch AI | |||
| LMCA | 16.6% | #102 of 125, top 82% | Epoch AI | ||
| Epoch Capabilities Index | 128.97 | #143 of 213, top 68% | Epoch AI | 2024-05-13 | |
| Epoch Capabilities Index | 128.76 | Epoch AI | 2024-08-06 | ||
| Epoch Capabilities Index | 128.81 | Epoch AI | 2024-11-20 | ||
| ForecastBench | 57.7 | #54 of 72, top 75% | Epoch AI | ||
| LiveBench | 52.2% | Epoch AI | |||
| LiveBench | 55.3% | #15 of 39, top 39% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 0.4% | #80 of 81, top 99% | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 6.3% | Epoch AI | 2025-02-25 | ||
| OTIS Mock AIME 2024-2025 | 6.3% | Epoch AI | 2025-02-25 | ||
| OTIS Mock AIME 2024-2025 | 6.4% | #144 of 173, top 84% | Epoch AI | 2025-02-25 | |
| Omni-MATH | 29.3% | #40 of 57, top 71% | HELM Capabilities | ||
| LiveBench Math | 42.9% | Epoch AI | |||
| LiveBench Math | 49.5% | #20 of 39, top 52% | Epoch AI | ||
| LMArena Math | 1284 | LMArena | 2026-10-08 | ||
| LMArena Math | 1285 | #189 of 285, top 67% | LMArena | 2026-10-08 | |
| MATH Level 5 | 49.8% | Epoch AI | 2025-02-05 | ||
| MATH Level 5 | 53.3% | #43 of 79, top 55% | Epoch AI | 2025-01-27 | |
| MATH Level 5 | 51% | Epoch AI | 2025-01-27 | ||
| FrontierMath (Feb 2025 set) | 0.3% | #65 of 68, top 96% | Epoch AI | 2025-03-07 | |
| FrontierMath (Feb 2025 set) | 0.3% | Epoch AI | 2025-03-06 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 48.9% | Epoch AI | 2025-01-27 | ||
| GPQA Diamond | 47.9% | Epoch AI | 2025-02-05 | ||
| GPQA Diamond | 49.2% | #128 of 186, top 69% | Epoch AI | 2025-01-27 | |
| Humanity's Last Exam | 2.7% | #41 of 41, top 100% | Epoch AI | ||
| SimpleQA Verified | 26% | #60 of 77, top 78% | Epoch AI | 2026-08-31 | |
| MMLU-Pro | 71.3% | #35 of 58, top 61% | HELM Capabilities | ||
| Confabulations (lower is better) | 15.3% | #18 of 51, top 36% | Lech Mazur benchmarks | ||
| Confabulations (lower is better) | 17.2% | Lech Mazur benchmarks | |||
| Vectara Hallucination Rate (lower is better) | 9.6% | #51 of 96, top 54% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 52% | #31 of 57, top 55% | HELM Capabilities | ||
| LMArena Expert | 1241 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1250 | #196 of 273, top 72% | LMArena | 2026-10-08 | |
| MMLU | 88.1% | Best of 81 | Epoch AI | ||
| MMLU | 84.2% | Epoch AI | |||
| MMLU | 84.3% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1137 | #102 of 122, top 84% | LMArena | 2026-10-09 | |
| LMArena Vision | 1065 | LMArena | 2026-10-09 | ||
| Video-MME | 71.9% | #6 of 15, top 40% | Epoch AI | ||
| Video-MME | 71.9% | #6 of 15, top 40% | Epoch AI | ||
| GeoBench | 71% | #12 of 25, top 48% | Epoch AI | ||
| VPCT | 40% | #13 of 24, top 55% | Epoch AI | ||
| ScienceQA | 88.5% | Best of 6 | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1262 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1283 | #186 of 297, top 63% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1254 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1277 | #193 of 285, top 68% | LMArena | 2026-10-08 | |
| LMArena French | 1304 | #160 of 223, top 72% | LMArena | 2026-10-08 | |
| LMArena French | 1263 | LMArena | 2026-10-08 | ||
| LMArena German | 1257 | LMArena | 2026-10-08 | ||
| LMArena German | 1282 | #157 of 231, top 68% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1234 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1257 | #138 of 211, top 66% | LMArena | 2026-10-08 | |
| LMArena Korean | 1218 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1234 | #150 of 213, top 71% | LMArena | 2026-10-08 | |
| LMArena Russian | 1272 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1286 | #186 of 283, top 66% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1292 | #163 of 226, top 73% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1269 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 64.9% | Epoch AI | |||
| LiveBench Instruction Following | 68.6% | #20 of 39, top 52% | Epoch AI | ||
| IFEval | 81.7% | #35 of 57, top 62% | HELM Capabilities | ||
| LMArena Instruction Following | 1278 | #190 of 298, top 64% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1267 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 66.7% | #19 of 47, top 41% | Epoch AI | ||
| LMArena Longer Query | 1289 | #198 of 291, top 69% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1283 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1300 | #190 of 297, top 64% | LMArena | 2026-10-08 | |
| LMArena Text | 1283 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1275 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1292 | #173 of 295, top 59% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 81.8% | #11 of 39, top 29% | Epoch AI | ||
| WildBench | 82.8% | #19 of 57, top 34% | HELM Capabilities | ||
| LMArena Multi-Turn | 1279 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1302 | #186 of 295, top 64% | LMArena | 2026-10-08 | |
| LiveBench Language | 47.4% | Epoch AI | |||
| LiveBench Language | 47.6% | #14 of 39, top 36% | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $2.50 | $10 | $1.25 | 2026-10-10 |
| openai | $2.50 | $10 | $1.25 | 2026-10-10 |
| openrouter | $2.50 | $10 | $1.25 | 2026-10-10 |
Compare GPT-4o
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-4o?
GPT-4o by OpenAI ranks 324th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.6. Its strongest category is multimodal, where it ranks 91st. API pricing starts at $2.50 per million input tokens and $10 per million output tokens, with a 128K-token context window.
How much does GPT-4o cost?
GPT-4o costs $2.50 per million input tokens and $10 per million output tokens on OpenAI's own API, with cached input at $1.25.
What is GPT-4o's context window?
GPT-4o accepts up to 128K tokens of input and can write up to 16K tokens in one response.
Is GPT-4o open source?
No. GPT-4o is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-4o's strengths and weaknesses?
Relative to other ranked models, GPT-4o places best in writing & preference, long context, multilingual and lowest in reasoning, coding, math.
What is GPT-4o best at?
Its best category is multimodal, where it ranks 91st on Noometry.