OpenAI, proprietary

GPT-4 Turbo

GPT-4 Turbo by OpenAI ranks 292nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.5. Its strongest category is multimodal, where it ranks 110th. API pricing starts at $10 per million input tokens and $30 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#292 of 354
Index score
30.5
Evidence
Confirmed 36 results
Provider
OpenAI
Released
November 6, 2023
Weights
Proprietary
Reasoning
No
Context window
128K
Max output
4K
Input price
$10 / M
Output price
$30 / M
Blended price
$15 / M
Output speed
Not measured
Value
#209 of 219
Knowledge cutoff
December 2023
Input
text, image

Category scores

Each category score combines every public result we have in that category.

GPT-4 Turbo category scores
  1. Coding 33.8
  2. Reasoning 15.3
  3. Math 9.0
  4. Knowledge 24.3
  5. Multimodal 30.6
  6. Multilingual 40.5
  7. Instruction Following 65.8
  8. Long Context 38.0
  9. Writing & Preference 47.7
GPT-4 Turbo category ranks
CategoryScoreRankResults
Coding33.8#2494
Reasoning15.3#3175
Math9.0#3224
Knowledge24.3#2683
Multimodal30.6#1101
Multilingual40.5#2161
Instruction Following65.8#2161
Long Context38.0#2061
Writing & Preference47.7#2063

Strengths and weaknesses

Categories where GPT-4 Turbo places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-4 Turbo: strongest categories
CategoryScorevs medianRank
Writing & Preference47.7−6.1#206 of 312, top 67%
Long Context38.0−2.9#206 of 296, top 70%
Instruction Following65.8−5.5#216 of 305, top 71%

Weakest categories

GPT-4 Turbo: weakest categories
CategoryScorevs medianRank
Math9.0−27.6#322 of 327, top 99%
Reasoning15.3−8.3#317 of 350, top 91%
Multimodal30.6−7.9#110 of 128, top 86%

Closest competitors

The models ranked just above and below GPT-4 Turbo. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4 Turbo
ModelRankScoreBlended $/MSpeed
Llama 3.1-405B#28830.7—78Compare
Yi-1.5-34B#28930.6——Compare
Codestral#29030.6$0.45271Compare
Llama-3.3-70B-Instruct#29130.6$0.16—Compare
Qwen1.5-32B#29330.5——Compare
Amazon Nova Micro#29430.4$0.0613—Compare
Olmo 7b Instruct#29530.3——Compare
Magistral Small#29630.2$0.750Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-4 Turbo Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML18%#107 of 119, top 90%Epoch AI
BigCodeBench Instruct48.2%#9 of 64, top 15%BigCodeBench2024-04-09
LMArena Coding1268#218 of 294, top 75%LMArena2026-10-08
BigCodeBench Complete58.2%#9 of 66, top 14%BigCodeBench2024-04-09
HumanEval+86.6%#6 of 45, top 14%april 2024EvalPlus
MBPP+73.3%#9 of 38, top 24%nov 2023EvalPlus

Agentic & Tool Use

GPT-4 Turbo Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
METR Time Horizons36.7%#27 of 32, top 85%Epoch AI
METR Time Horizons28.9%Epoch AI
METR Time Horizons35.2%Epoch AI

Reasoning

GPT-4 Turbo Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench25.1%#66 of 77, top 86%Epoch AI
Chess Puzzles6%#91 of 129, top 71%Epoch AI2026-07-15
LMArena Hard Prompts1251#221 of 297, top 75%LMArena2026-10-08
DTBench61.6%#112 of 151, top 75%Epoch AI
LMCA9.8%#112 of 125, top 90%Epoch AI
Epoch Capabilities Index126.46Epoch AI2024-01-25
Epoch Capabilities Index127.25#149 of 213, top 70%Epoch AI2024-04-09
ForecastBench59.4#39 of 72, top 55%Epoch AI

Math

GPT-4 Turbo Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)0.7%#78 of 81, top 97%Epoch AI2026-08-27
OTIS Mock AIME 2024-20256.7%#142 of 173, top 83%Epoch AI2025-02-27
LMArena Math1272#200 of 285, top 71%LMArena2026-10-08
MATH Level 535.4%Epoch AI2025-01-27
MATH Level 540%Epoch AI2025-01-27
MATH Level 546.7%#48 of 79, top 61%Epoch AI2025-02-27

Knowledge

GPT-4 Turbo Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond46.6%#139 of 186, top 75%Epoch AI2025-02-27
GPQA Diamond42.4%Epoch AI2025-01-27
GPQA Diamond42.3%Epoch AI2025-01-27
Confabulations (lower is better)28.4%#45 of 51, top 89%Lech Mazur benchmarks
LMArena Expert1223#211 of 273, top 78%LMArena2026-10-08
MMLU81.3%#14 of 81, top 18%Epoch AI
MMLU79.6%Epoch AI

Multimodal

GPT-4 Turbo Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1090#109 of 122, top 90%LMArena2026-10-09

Multilingual

GPT-4 Turbo Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1245#216 of 297, top 73%LMArena2026-10-08
LMArena Chinese1242#212 of 285, top 75%LMArena2026-10-08
LMArena French1276#173 of 223, top 78%LMArena2026-10-08
LMArena German1259#170 of 231, top 74%LMArena2026-10-08
LMArena Japanese1194#163 of 211, top 78%LMArena2026-10-08
LMArena Korean1187#169 of 213, top 80%LMArena2026-10-08
LMArena Russian1259#205 of 283, top 73%LMArena2026-10-08
LMArena Spanish1260#178 of 226, top 79%LMArena2026-10-08

Instruction Following

GPT-4 Turbo Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1249#212 of 298, top 72%LMArena2026-10-08

Long Context

GPT-4 Turbo Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1254#219 of 291, top 76%LMArena2026-10-08

Writing & Preference

GPT-4 Turbo Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1272#215 of 297, top 73%LMArena2026-10-08
LMArena Creative Writing1269#192 of 295, top 66%LMArena2026-10-08
LMArena Multi-Turn1267#214 of 295, top 73%LMArena2026-10-08

API pricing by provider

GPT-4 Turbo API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$10$30—2026-10-10
openai$10$30—2026-10-10
openrouter$10$30—2026-10-10

Compare GPT-4 Turbo

Other OpenAI models

Frequently asked questions

How good is GPT-4 Turbo?

GPT-4 Turbo by OpenAI ranks 292nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.5. Its strongest category is multimodal, where it ranks 110th. API pricing starts at $10 per million input tokens and $30 per million output tokens, with a 128K-token context window.

How much does GPT-4 Turbo cost?

GPT-4 Turbo costs $10 per million input tokens and $30 per million output tokens on OpenAI's own API.

What is GPT-4 Turbo's context window?

GPT-4 Turbo accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is GPT-4 Turbo open source?

No. GPT-4 Turbo is proprietary and available only through OpenAI's API and partner platforms.

What are GPT-4 Turbo's strengths and weaknesses?

Relative to other ranked models, GPT-4 Turbo places best in writing & preference, long context, instruction following and lowest in math, reasoning, multimodal.

What is GPT-4 Turbo best at?

Its best category is multimodal, where it ranks 110th on Noometry.