OpenAI, proprietary

GPT-4

GPT-4 by OpenAI ranks 316th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.1. Its strongest category is long context, where it ranks 212th. API pricing starts at $30 per million input tokens and $60 per million output tokens, with a 8K-token context window.

Last verified

Specifications

Noometry rank
#316 of 354
Index score
29.1
Evidence
Confirmed 38 results
Provider
OpenAI
Released
March 14, 2023
Weights
Proprietary
Reasoning
No
Context window
8K
Max output
8K
Input price
$30 / M
Output price
$60 / M
Blended price
$37.50 / M
Output speed
Not measured
Value
#218 of 219
Knowledge cutoff
November 2023
Input
text

Category scores

Each category score combines every public result we have in that category.

GPT-4 category scores
  1. Coding 31.6
  2. Reasoning 17.8
  3. Math 10.8
  4. Knowledge 18.4
  5. Multilingual 40.6
  6. Instruction Following 65.3
  7. Long Context 37.7
  8. Writing & Preference 34.9
GPT-4 category ranks
CategoryScoreRankResults
Coding31.6#2834
Reasoning17.8#2895
Math10.8#3093
Knowledge18.4#2822
Multilingual40.6#2151
Instruction Following65.3#2221
Long Context37.7#2121
Writing & Preference34.9#2684

Strengths and weaknesses

Categories where GPT-4 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-4: strongest categories
CategoryScorevs medianRank
Long Context37.7−3.2#212 of 296, top 72%
Multilingual40.6−6.8#215 of 297, top 73%
Instruction Following65.3−6.0#222 of 305, top 73%

Weakest categories

GPT-4: weakest categories
CategoryScorevs medianRank
Math10.8−25.8#309 of 327, top 95%
Knowledge18.4−18.9#282 of 314, top 90%
Writing & Preference34.9−18.9#268 of 312, top 86%

Closest competitors

The models ranked just above and below GPT-4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4
ModelRankScoreBlended $/MSpeed
Gemma 2 27B#31229.4$0.65—Compare
Gemma 1.1 2b IT#31329.3——Compare
Phi 3 Small 8k Instruct#31429.3——Compare
Claude 3.5 Haiku#31529.2——Compare
Llama 2-7B#31729.1——Compare
Granite 4.0 Micro#31829.0$0.0408—Compare
Claude 3 Sonnet#31929.0——Compare
Qwen2.5 7B Instruct#32029.0$0.31—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-4 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML12.4%#110 of 119, top 93%Epoch AI
BigCodeBench Instruct46%#15 of 64, top 24%BigCodeBench2024-06-13
LMArena Coding1249LMArena2026-10-08
LMArena Coding1187LMArena2026-10-08
LMArena Coding1254#224 of 294, top 77%LMArena2026-10-08
LMArena Coding1209LMArena2026-10-08
BigCodeBench Complete57.2%#15 of 66, top 23%BigCodeBench2024-06-13
HumanEval+79.3%#12 of 45, top 27%may 2023EvalPlus

Agentic & Tool Use

GPT-4 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
METR Time Horizons36.1%#28 of 32, top 88%Epoch AI

Reasoning

GPT-4 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles4%#97 of 129, top 76%Epoch AI2026-08-07
LMArena Hard Prompts1241#225 of 297, top 76%LMArena2026-10-08
LMArena Hard Prompts1174LMArena2026-10-08
LMArena Hard Prompts1200LMArena2026-10-08
LMArena Hard Prompts1239LMArena2026-10-08
Mystery Game Puzzles12%#56 of 74, top 76%Epoch AI2026-08-28
DTBench62.7%#108 of 151, top 72%Epoch AI
LMCA17.1%#99 of 125, top 80%Epoch AI
BIG-Bench Hard75.1%#8 of 27, top 30%Epoch AI
Epoch Capabilities Index125.89#155 of 213, top 73%Epoch AI2023-03-14
Epoch Capabilities Index123.12Epoch AI2023-06-13
ForecastBench57.8#53 of 72, top 74%Epoch AI
HellaSwag95.3%Best of 29Epoch AI
HellaSwag95.3%Best of 29Epoch AI
WinoGrande87.5%#3 of 43, top 7%Epoch AI
WinoGrande87.5%#3 of 43, top 7%Epoch AI

Math

GPT-4 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20250.6%Epoch AI2025-10-22
OTIS Mock AIME 2024-20251.1%#168 of 173, top 98%Epoch AI2025-10-23
LMArena Math1230LMArena2026-10-08
LMArena Math1217LMArena2026-10-08
LMArena Math1268LMArena2026-10-08
LMArena Math1269#202 of 285, top 71%LMArena2026-10-08
MATH Level 523%#60 of 79, top 76%Epoch AI2025-01-27
GSM8K92%#3 of 38, top 8%Epoch AI
GSM8K90%Epoch AI

Knowledge

GPT-4 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond30.7%Epoch AI2025-01-27
GPQA Diamond35.7%#157 of 186, top 85%Epoch AI2025-10-23
LMArena Expert1128LMArena2026-10-08
LMArena Expert1149LMArena2026-10-08
LMArena Expert1205LMArena2026-10-08
LMArena Expert1211#216 of 273, top 80%LMArena2026-10-08
MMLU86.4%#5 of 81, top 7%Epoch AI
MMLU82.4%Epoch AI
TriviaQA84.8%#5 of 25, top 20%Epoch AI

Multilingual

GPT-4 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1241LMArena2026-10-08
LMArena Non-English1194LMArena2026-10-08
LMArena Non-English1160LMArena2026-10-08
LMArena Non-English1246#215 of 297, top 73%LMArena2026-10-08
LMArena Chinese1238LMArena2026-10-08
LMArena Chinese1184LMArena2026-10-08
LMArena Chinese1136LMArena2026-10-08
LMArena Chinese1242#211 of 285, top 75%LMArena2026-10-08
LMArena French1281LMArena2026-10-08
LMArena French1170LMArena2026-10-08
LMArena French1219LMArena2026-10-08
LMArena French1283#168 of 223, top 76%LMArena2026-10-08
LMArena German1198LMArena2026-10-08
LMArena German1251#178 of 231, top 78%LMArena2026-10-08
LMArena German1162LMArena2026-10-08
LMArena German1246LMArena2026-10-08
LMArena Japanese1209#152 of 211, top 73%LMArena2026-10-08
LMArena Japanese1195LMArena2026-10-08
LMArena Japanese1114LMArena2026-10-08
LMArena Japanese1137LMArena2026-10-08
LMArena Korean1184#170 of 213, top 80%LMArena2026-10-08
LMArena Korean1174LMArena2026-10-08
LMArena Korean1087LMArena2026-10-08
LMArena Korean1060LMArena2026-10-08
LMArena Russian1251#213 of 283, top 76%LMArena2026-10-08
LMArena Russian1173LMArena2026-10-08
LMArena Russian1241LMArena2026-10-08
LMArena Russian1195LMArena2026-10-08
LMArena Spanish1168LMArena2026-10-08
LMArena Spanish1261#177 of 226, top 79%LMArena2026-10-08
LMArena Spanish1197LMArena2026-10-08
LMArena Spanish1248LMArena2026-10-08

Instruction Following

GPT-4 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1235LMArena2026-10-08
LMArena Instruction Following1201LMArena2026-10-08
LMArena Instruction Following1189LMArena2026-10-08
LMArena Instruction Following1241#221 of 298, top 75%LMArena2026-10-08

Long Context

GPT-4 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1190LMArena2026-10-08
LMArena Longer Query1188LMArena2026-10-08
LMArena Longer Query1244#224 of 291, top 77%LMArena2026-10-08
LMArena Longer Query1236LMArena2026-10-08

Writing & Preference

GPT-4 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1206LMArena2026-10-08
LMArena Text1262LMArena2026-10-08
LMArena Text1186LMArena2026-10-08
LMArena Text1263#219 of 297, top 74%LMArena2026-10-08
LMArena Creative Writing1232LMArena2026-10-08
LMArena Creative Writing1190LMArena2026-10-08
LMArena Creative Writing1244#210 of 295, top 72%LMArena2026-10-08
LMArena Creative Writing1192LMArena2026-10-08
EQ-Bench Creative Writing752#107 of 115, top 94%EQ-Bench
LMArena Multi-Turn1257#218 of 295, top 74%LMArena2026-10-08
LMArena Multi-Turn1185LMArena2026-10-08
LMArena Multi-Turn1206LMArena2026-10-08
LMArena Multi-Turn1250LMArena2026-10-08

API pricing by provider

GPT-4 API prices
RouteInput $/MOutput $/MCached input $/MChecked
openai$30$60—2026-10-10
openrouter$30$60—2026-10-10

Compare GPT-4

Other OpenAI models

Frequently asked questions

How good is GPT-4?

GPT-4 by OpenAI ranks 316th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.1. Its strongest category is long context, where it ranks 212th. API pricing starts at $30 per million input tokens and $60 per million output tokens, with a 8K-token context window.

How much does GPT-4 cost?

GPT-4 costs $30 per million input tokens and $60 per million output tokens on OpenAI's own API.

What is GPT-4's context window?

GPT-4 accepts up to 8K tokens of input and can write up to 8K tokens in one response.

Is GPT-4 open source?

No. GPT-4 is proprietary and available only through OpenAI's API and partner platforms.

What are GPT-4's strengths and weaknesses?

Relative to other ranked models, GPT-4 places best in long context, multilingual, instruction following and lowest in math, knowledge, writing & preference.

What is GPT-4 best at?

Its best category is long context, where it ranks 212th on Noometry.