OpenAI, proprietary
GPT-4
GPT-4 by OpenAI ranks 316th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.1. Its strongest category is long context, where it ranks 212th. API pricing starts at $30 per million input tokens and $60 per million output tokens, with a 8K-token context window.
Last verified
Specifications
- Noometry rank
- #316 of 354
- Index score
- 29.1
- Evidence
- Confirmed 38 results
- Provider
- OpenAI
- Released
- March 14, 2023
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 8K
- Max output
- 8K
- Input price
- $30 / M
- Output price
- $60 / M
- Blended price
- $37.50 / M
- Output speed
- Not measured
- Value
- #218 of 219
- Knowledge cutoff
- November 2023
- Input
- text
Category scores
Each category score combines every public result we have in that category.
- Coding 31.6
- Reasoning 17.8
- Math 10.8
- Knowledge 18.4
- Multilingual 40.6
- Instruction Following 65.3
- Long Context 37.7
- Writing & Preference 34.9
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 31.6 | #283 | 4 |
| Reasoning | 17.8 | #289 | 5 |
| Math | 10.8 | #309 | 3 |
| Knowledge | 18.4 | #282 | 2 |
| Multilingual | 40.6 | #215 | 1 |
| Instruction Following | 65.3 | #222 | 1 |
| Long Context | 37.7 | #212 | 1 |
| Writing & Preference | 34.9 | #268 | 4 |
Strengths and weaknesses
Categories where GPT-4 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 37.7 | −3.2 | #212 of 296, top 72% |
| Multilingual | 40.6 | −6.8 | #215 of 297, top 73% |
| Instruction Following | 65.3 | −6.0 | #222 of 305, top 73% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 10.8 | −25.8 | #309 of 327, top 95% |
| Knowledge | 18.4 | −18.9 | #282 of 314, top 90% |
| Writing & Preference | 34.9 | −18.9 | #268 of 312, top 86% |
Closest competitors
The models ranked just above and below GPT-4. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemma 2 27B | #312 | 29.4 | $0.65 | — | Compare |
| Gemma 1.1 2b IT | #313 | 29.3 | — | — | Compare |
| Phi 3 Small 8k Instruct | #314 | 29.3 | — | — | Compare |
| Claude 3.5 Haiku | #315 | 29.2 | — | — | Compare |
| Llama 2-7B | #317 | 29.1 | — | — | Compare |
| Granite 4.0 Micro | #318 | 29.0 | $0.0408 | — | Compare |
| Claude 3 Sonnet | #319 | 29.0 | — | — | Compare |
| Qwen2.5 7B Instruct | #320 | 29.0 | $0.31 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 12.4% | #110 of 119, top 93% | Epoch AI | ||
| BigCodeBench Instruct | 46% | #15 of 64, top 24% | BigCodeBench | 2024-06-13 | |
| LMArena Coding | 1249 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1187 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1254 | #224 of 294, top 77% | LMArena | 2026-10-08 | |
| LMArena Coding | 1209 | LMArena | 2026-10-08 | ||
| BigCodeBench Complete | 57.2% | #15 of 66, top 23% | BigCodeBench | 2024-06-13 | |
| HumanEval+ | 79.3% | #12 of 45, top 27% | may 2023 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| METR Time Horizons | 36.1% | #28 of 32, top 88% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 4% | #97 of 129, top 76% | Epoch AI | 2026-08-07 | |
| LMArena Hard Prompts | 1241 | #225 of 297, top 76% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1174 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1200 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1239 | LMArena | 2026-10-08 | ||
| Mystery Game Puzzles | 12% | #56 of 74, top 76% | Epoch AI | 2026-08-28 | |
| DTBench | 62.7% | #108 of 151, top 72% | Epoch AI | ||
| LMCA | 17.1% | #99 of 125, top 80% | Epoch AI | ||
| BIG-Bench Hard | 75.1% | #8 of 27, top 30% | Epoch AI | ||
| Epoch Capabilities Index | 125.89 | #155 of 213, top 73% | Epoch AI | 2023-03-14 | |
| Epoch Capabilities Index | 123.12 | Epoch AI | 2023-06-13 | ||
| ForecastBench | 57.8 | #53 of 72, top 74% | Epoch AI | ||
| HellaSwag | 95.3% | Best of 29 | Epoch AI | ||
| HellaSwag | 95.3% | Best of 29 | Epoch AI | ||
| WinoGrande | 87.5% | #3 of 43, top 7% | Epoch AI | ||
| WinoGrande | 87.5% | #3 of 43, top 7% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 0.6% | Epoch AI | 2025-10-22 | ||
| OTIS Mock AIME 2024-2025 | 1.1% | #168 of 173, top 98% | Epoch AI | 2025-10-23 | |
| LMArena Math | 1230 | LMArena | 2026-10-08 | ||
| LMArena Math | 1217 | LMArena | 2026-10-08 | ||
| LMArena Math | 1268 | LMArena | 2026-10-08 | ||
| LMArena Math | 1269 | #202 of 285, top 71% | LMArena | 2026-10-08 | |
| MATH Level 5 | 23% | #60 of 79, top 76% | Epoch AI | 2025-01-27 | |
| GSM8K | 92% | #3 of 38, top 8% | Epoch AI | ||
| GSM8K | 90% | Epoch AI |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 30.7% | Epoch AI | 2025-01-27 | ||
| GPQA Diamond | 35.7% | #157 of 186, top 85% | Epoch AI | 2025-10-23 | |
| LMArena Expert | 1128 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1149 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1205 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1211 | #216 of 273, top 80% | LMArena | 2026-10-08 | |
| MMLU | 86.4% | #5 of 81, top 7% | Epoch AI | ||
| MMLU | 82.4% | Epoch AI | |||
| TriviaQA | 84.8% | #5 of 25, top 20% | Epoch AI |
Multilingual
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1235 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1201 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1189 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1241 | #221 of 298, top 75% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1190 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1188 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1244 | #224 of 291, top 77% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1236 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1206 | LMArena | 2026-10-08 | ||
| LMArena Text | 1262 | LMArena | 2026-10-08 | ||
| LMArena Text | 1186 | LMArena | 2026-10-08 | ||
| LMArena Text | 1263 | #219 of 297, top 74% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1232 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1190 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1244 | #210 of 295, top 72% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1192 | LMArena | 2026-10-08 | ||
| EQ-Bench Creative Writing | 752 | #107 of 115, top 94% | EQ-Bench | ||
| LMArena Multi-Turn | 1257 | #218 of 295, top 74% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1185 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1206 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1250 | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| openai | $30 | $60 | — | 2026-10-10 |
| openrouter | $30 | $60 | — | 2026-10-10 |
Compare GPT-4
- GPT-4 vs Claude 3.5 Haiku
- GPT-4 vs Llama 2-7B
- GPT-4 vs Phi 3 Small 8k Instruct
- GPT-4 vs Granite 4.0 Micro
- GPT-4 vs Gemma 1.1 2b IT
- GPT-4 vs Claude 3 Sonnet
- GPT-4 vs Claude Fable 5.1
- GPT-4 vs Gemini 3.8 Flash
- GPT-4 vs Kimi K3
- GPT-4 vs Grok 4.6
- GPT-4 vs Qwen3.8 Max
- GPT-4 vs GLM-5.3
- GPT-4 vs Muse Spark 1.3
- GPT-4 vs DeepSeek V4 Pro
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-4?
GPT-4 by OpenAI ranks 316th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.1. Its strongest category is long context, where it ranks 212th. API pricing starts at $30 per million input tokens and $60 per million output tokens, with a 8K-token context window.
How much does GPT-4 cost?
GPT-4 costs $30 per million input tokens and $60 per million output tokens on OpenAI's own API.
What is GPT-4's context window?
GPT-4 accepts up to 8K tokens of input and can write up to 8K tokens in one response.
Is GPT-4 open source?
No. GPT-4 is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-4's strengths and weaknesses?
Relative to other ranked models, GPT-4 places best in long context, multilingual, instruction following and lowest in math, knowledge, writing & preference.
What is GPT-4 best at?
Its best category is long context, where it ranks 212th on Noometry.