OpenAI, proprietary
GPT-4 Turbo
GPT-4 Turbo by OpenAI ranks 292nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.5. Its strongest category is multimodal, where it ranks 110th. API pricing starts at $10 per million input tokens and $30 per million output tokens, with a 128K-token context window.
Last verified
Specifications
- Noometry rank
- #292 of 354
- Index score
- 30.5
- Evidence
- Confirmed 36 results
- Provider
- OpenAI
- Released
- November 6, 2023
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 128K
- Max output
- 4K
- Input price
- $10 / M
- Output price
- $30 / M
- Blended price
- $15 / M
- Output speed
- Not measured
- Value
- #209 of 219
- Knowledge cutoff
- December 2023
- Input
- text, image
Category scores
Each category score combines every public result we have in that category.
- Coding 33.8
- Reasoning 15.3
- Math 9.0
- Knowledge 24.3
- Multimodal 30.6
- Multilingual 40.5
- Instruction Following 65.8
- Long Context 38.0
- Writing & Preference 47.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 33.8 | #249 | 4 |
| Reasoning | 15.3 | #317 | 5 |
| Math | 9.0 | #322 | 4 |
| Knowledge | 24.3 | #268 | 3 |
| Multimodal | 30.6 | #110 | 1 |
| Multilingual | 40.5 | #216 | 1 |
| Instruction Following | 65.8 | #216 | 1 |
| Long Context | 38.0 | #206 | 1 |
| Writing & Preference | 47.7 | #206 | 3 |
Strengths and weaknesses
Categories where GPT-4 Turbo places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 47.7 | −6.1 | #206 of 312, top 67% |
| Long Context | 38.0 | −2.9 | #206 of 296, top 70% |
| Instruction Following | 65.8 | −5.5 | #216 of 305, top 71% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 9.0 | −27.6 | #322 of 327, top 99% |
| Reasoning | 15.3 | −8.3 | #317 of 350, top 91% |
| Multimodal | 30.6 | −7.9 | #110 of 128, top 86% |
Closest competitors
The models ranked just above and below GPT-4 Turbo. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Llama 3.1-405B | #288 | 30.7 | — | 78 | Compare |
| Yi-1.5-34B | #289 | 30.6 | — | — | Compare |
| Codestral | #290 | 30.6 | $0.45 | 271 | Compare |
| Llama-3.3-70B-Instruct | #291 | 30.6 | $0.16 | — | Compare |
| Qwen1.5-32B | #293 | 30.5 | — | — | Compare |
| Amazon Nova Micro | #294 | 30.4 | $0.0613 | — | Compare |
| Olmo 7b Instruct | #295 | 30.3 | — | — | Compare |
| Magistral Small | #296 | 30.2 | $0.75 | 0 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 18% | #107 of 119, top 90% | Epoch AI | ||
| BigCodeBench Instruct | 48.2% | #9 of 64, top 15% | BigCodeBench | 2024-04-09 | |
| LMArena Coding | 1268 | #218 of 294, top 75% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 58.2% | #9 of 66, top 14% | BigCodeBench | 2024-04-09 | |
| HumanEval+ | 86.6% | #6 of 45, top 14% | april 2024 | EvalPlus | |
| MBPP+ | 73.3% | #9 of 38, top 24% | nov 2023 | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| METR Time Horizons | 36.7% | #27 of 32, top 85% | Epoch AI | ||
| METR Time Horizons | 28.9% | Epoch AI | |||
| METR Time Horizons | 35.2% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 25.1% | #66 of 77, top 86% | Epoch AI | ||
| Chess Puzzles | 6% | #91 of 129, top 71% | Epoch AI | 2026-07-15 | |
| LMArena Hard Prompts | 1251 | #221 of 297, top 75% | LMArena | 2026-10-08 | |
| DTBench | 61.6% | #112 of 151, top 75% | Epoch AI | ||
| LMCA | 9.8% | #112 of 125, top 90% | Epoch AI | ||
| Epoch Capabilities Index | 126.46 | Epoch AI | 2024-01-25 | ||
| Epoch Capabilities Index | 127.25 | #149 of 213, top 70% | Epoch AI | 2024-04-09 | |
| ForecastBench | 59.4 | #39 of 72, top 55% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 0.7% | #78 of 81, top 97% | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 6.7% | #142 of 173, top 83% | Epoch AI | 2025-02-27 | |
| LMArena Math | 1272 | #200 of 285, top 71% | LMArena | 2026-10-08 | |
| MATH Level 5 | 35.4% | Epoch AI | 2025-01-27 | ||
| MATH Level 5 | 40% | Epoch AI | 2025-01-27 | ||
| MATH Level 5 | 46.7% | #48 of 79, top 61% | Epoch AI | 2025-02-27 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 46.6% | #139 of 186, top 75% | Epoch AI | 2025-02-27 | |
| GPQA Diamond | 42.4% | Epoch AI | 2025-01-27 | ||
| GPQA Diamond | 42.3% | Epoch AI | 2025-01-27 | ||
| Confabulations (lower is better) | 28.4% | #45 of 51, top 89% | Lech Mazur benchmarks | ||
| LMArena Expert | 1223 | #211 of 273, top 78% | LMArena | 2026-10-08 | |
| MMLU | 81.3% | #14 of 81, top 18% | Epoch AI | ||
| MMLU | 79.6% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1090 | #109 of 122, top 90% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1245 | #216 of 297, top 73% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1242 | #212 of 285, top 75% | LMArena | 2026-10-08 | |
| LMArena French | 1276 | #173 of 223, top 78% | LMArena | 2026-10-08 | |
| LMArena German | 1259 | #170 of 231, top 74% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1194 | #163 of 211, top 78% | LMArena | 2026-10-08 | |
| LMArena Korean | 1187 | #169 of 213, top 80% | LMArena | 2026-10-08 | |
| LMArena Russian | 1259 | #205 of 283, top 73% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1260 | #178 of 226, top 79% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1249 | #212 of 298, top 72% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1254 | #219 of 291, top 76% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1272 | #215 of 297, top 73% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1269 | #192 of 295, top 66% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1267 | #214 of 295, top 73% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $10 | $30 | — | 2026-10-10 |
| openai | $10 | $30 | — | 2026-10-10 |
| openrouter | $10 | $30 | — | 2026-10-10 |
Compare GPT-4 Turbo
- GPT-4 Turbo vs GPT-3.5-turbo
- GPT-4 Turbo vs Llama-3.3-70B-Instruct
- GPT-4 Turbo vs Qwen1.5-32B
- GPT-4 Turbo vs Codestral
- GPT-4 Turbo vs Amazon Nova Micro
- GPT-4 Turbo vs Yi-1.5-34B
- GPT-4 Turbo vs Olmo 7b Instruct
- GPT-4 Turbo vs Claude Fable 5.1
- GPT-4 Turbo vs Gemini 3.8 Flash
- GPT-4 Turbo vs Kimi K3
- GPT-4 Turbo vs Grok 4.6
- GPT-4 Turbo vs Qwen3.8 Max
- GPT-4 Turbo vs GLM-5.3
- GPT-4 Turbo vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-4 Turbo?
GPT-4 Turbo by OpenAI ranks 292nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.5. Its strongest category is multimodal, where it ranks 110th. API pricing starts at $10 per million input tokens and $30 per million output tokens, with a 128K-token context window.
How much does GPT-4 Turbo cost?
GPT-4 Turbo costs $10 per million input tokens and $30 per million output tokens on OpenAI's own API.
What is GPT-4 Turbo's context window?
GPT-4 Turbo accepts up to 128K tokens of input and can write up to 4K tokens in one response.
Is GPT-4 Turbo open source?
No. GPT-4 Turbo is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-4 Turbo's strengths and weaknesses?
Relative to other ranked models, GPT-4 Turbo places best in writing & preference, long context, instruction following and lowest in math, reasoning, multimodal.
What is GPT-4 Turbo best at?
Its best category is multimodal, where it ranks 110th on Noometry.