OpenAI, proprietary

GPT-5

GPT-5 by OpenAI ranks 45th of 354 ranked models on the Noometry Index as of October 2026, with a score of 50.9. Its strongest category is long context, where it ranks 2nd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

Last verified

Specifications

Noometry rank
#45 of 354
Index score
50.9
Evidence
Confirmed 69 results
Provider
OpenAI
Released
August 7, 2025
Weights
Proprietary
Reasoning
Yes
Context window
400K
Max output
128K
Input price
$1.25 / M
Output price
$10 / M
Blended price
$3.44 / M
Output speed
2 tokens/s Kagi
Value
#164 of 219
Knowledge cutoff
September 2024
Input
text, image

Category scores

Each category score combines every public result we have in that category.

GPT-5 category scores
  1. Coding 50.3
  2. Agentic & Tool Use 33.1
  3. Reasoning 38.3
  4. Math 55.0
  5. Knowledge 56.6
  6. Multimodal 46.8
  7. Multilingual 51.4
  8. Instruction Following 73.8
  9. Long Context 69.5
  10. Writing & Preference 63.4
GPT-5 category ranks
CategoryScoreRankResults
Coding50.3#478
Agentic & Tool Use33.1#565
Reasoning38.3#6412
Math55.0#447
Knowledge56.6#438
Multimodal46.8#133
Multilingual51.4#1101
Instruction Following73.8#1132
Long Context69.5#22
Writing & Preference63.4#656

Strengths and weaknesses

Categories where GPT-5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5: strongest categories
CategoryScorevs medianRank
Long Context69.5+28.6#2 of 296, top 1%
Multimodal46.8+8.3#13 of 128, top 11%
Math55.0+18.4#44 of 327, top 14%

Weakest categories

GPT-5: weakest categories
CategoryScorevs medianRank
Instruction Following73.8+2.6#113 of 305, top 38%
Multilingual51.4+3.9#110 of 297, top 38%
Agentic & Tool Use33.1+2.7#56 of 154, top 37%

Closest competitors

The models ranked just above and below GPT-5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5
ModelRankScoreBlended $/MSpeed
GLM-5.3-Flash#4151.8$0.24—Compare
Qwen3.7 Max#4251.5$3.75—Compare
Qwen3.6 Max Preview#4351.5$2.92—Compare
GLM-5.2#4451.1$2.1523Compare
Muse Spark#4650.6——Compare
Claude Opus 4.5#4750.5$1013Compare
Muse Spark 1.2#4850.3$2—Compare
MiMo-V2.6-Pro#4950.3$0.54—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified73.6%#19 of 32, top 60%highEpoch AI2026-02-06
SWE-bench Verified71.5%mediumEpoch AI2026-02-05
SWE-bench Verified (bash only)65%#16 of 39, top 42%mediumSWE-bench2025-08-07
Aider Polyglot88%Best of 44highEpoch AI
Aider Polyglot81.3%lowEpoch AI
Aider Polyglot86.7%mediumEpoch AI
LMArena WebDev1418#65 of 113, top 58%mediumLMArena2026-10-08
SciCode42.9%#67 of 121, top 56%Epoch AI
GSO6.9%#20 of 31, top 65%highEpoch AI
WeirdML39.8%Epoch AI
WeirdML60.7%#30 of 119, top 26%highEpoch AI
LMArena Coding1398LMArena2026-10-08
LMArena Coding1436#102 of 294, top 35%highLMArena2026-10-08
ALE-Bench1,162#34 of 105, top 33%highEpoch AI
ALE-Bench807.65minimalEpoch AI
AlgoTune1.67#8 of 18, top 45%highEpoch AI

Agentic & Tool Use

GPT-5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench49.6%#17 of 41, top 42%Epoch AI
Terminal-Bench49.6%mediumEpoch AI
GDPval34.8%#6 of 11, top 55%mediumEpoch AI
Remote Labor Index1.7%#12 of 14, top 86%Epoch AI
DeepResearch Bench48.1%highEpoch AI
DeepResearch Bench49.6%#8 of 24, top 34%lowEpoch AI
DeepResearch Bench48.6%mediumEpoch AI
DeepResearch Bench48.9%minimalEpoch AI
BALROG32.8%#14 of 35, top 40%minimalEpoch AI
LMArena Search1133#29 of 32, top 91%LMArena2026-08-24
METR Time Horizons69.4%highEpoch AI
METR Time Horizons69.6%#10 of 32, top 32%mediumEpoch AI

Reasoning

GPT-5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-29.9%#49 of 83, top 60%highEpoch AI
ARC-AGI-21.9%lowEpoch AI
ARC-AGI-27.5%mediumEpoch AI
ARC-AGI-20%minimalEpoch AI
SimpleBench56.7%#31 of 77, top 41%highEpoch AI
Kagi LLM Benchmark72.7%#19 of 99, top 20%Kagi LLM Benchmark
ARC-AGI-165.7%#45 of 83, top 55%highEpoch AI
ARC-AGI-144%lowEpoch AI
ARC-AGI-156.2%mediumEpoch AI
ARC-AGI-16%minimalEpoch AI
CritPt12.6%#42 of 134, top 32%highEpoch AI
CritPt0%minimalEpoch AI
Chess Puzzles37%#27 of 129, top 21%highEpoch AI2025-12-08
Chess Puzzles24%lowEpoch AI2026-07-15
Chess Puzzles29%mediumEpoch AI2026-07-15
Chess Puzzles16%minimalEpoch AI2026-07-15
EnigmaEval10.5%#13 of 38, top 35%Epoch AI
EBR-Bench12.7%#18 of 24, top 75%highEpoch AI2026-06-29
LMArena Hard Prompts1405LMArena2026-10-08
LMArena Hard Prompts1416#110 of 297, top 38%highLMArena2026-10-08
Mystery Game Puzzles23%#35 of 74, top 48%highEpoch AI2026-08-05
Mystery Game Puzzles16%lowEpoch AI2026-08-27
Mystery Game Puzzles15%mediumEpoch AI2026-08-27
Mystery Game Puzzles14%minimalEpoch AI2026-08-27
DTBench90.7%#35 of 151, top 24%highEpoch AI
LMCA40%#45 of 125, top 36%highEpoch AI
Epoch Capabilities Index150#53 of 213, top 25%Epoch AI2025-08-07
ForecastBench61.4#10 of 72, top 14%Epoch AI

Math

GPT-5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)55.4%#43 of 81, top 54%highEpoch AI2026-06-10
FrontierMath (Tiers 1-3)37.2%lowEpoch AI2026-08-27
FrontierMath (Tiers 1-3)18.2%minimalEpoch AI2026-08-27
FrontierMath Tier 422%#42 of 63, top 67%highEpoch AI2026-06-11
OTIS Mock AIME 2024-202591.4%#48 of 173, top 28%highEpoch AI2025-10-29
OTIS Mock AIME 2024-202587.2%mediumEpoch AI2025-08-07
OTIS Mock AIME 2024-202546.7%minimalEpoch AI2026-07-20
ProofBench18%#51 of 77, top 67%highEpoch AI
Omni-MATH64.7%#7 of 57, top 13%HELM Capabilities
LMArena Math1407#116 of 285, top 41%LMArena2026-10-08
LMArena Math1399highLMArena2026-10-08
MATH Level 598.1%Best of 79highEpoch AI2025-10-29
MATH Level 597.9%mediumEpoch AI2025-08-20
FrontierMath (Feb 2025 set)32.4%#16 of 68, top 24%highEpoch AI2025-11-13
FrontierMath (Feb 2025 set)27.2%mediumEpoch AI2025-11-13
FrontierMath Tier 4 (v1)12.5%#18 of 55, top 33%highEpoch AI2025-10-30
FrontierMath Tier 4 (v1)6.3%mediumEpoch AI2025-08-07

Knowledge

GPT-5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond86.2%#59 of 186, top 32%highEpoch AI2025-10-29
GPQA Diamond85.4%mediumEpoch AI2025-08-07
GPQA Diamond71.7%minimalEpoch AI2026-07-20
Humanity's Last Exam25.3%#13 of 41, top 32%Epoch AI
Humanity's Last Exam25.3%highEpoch AI
SimpleQA Verified50.1%#25 of 77, top 33%highEpoch AI2026-08-27
MMLU-Pro86.3%#5 of 58, top 9%HELM Capabilities
Confabulations (lower is better)10.3%Best of 51medium reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)15.1%Vectara Hallucination Leaderboard
Vectara Hallucination Rate (lower is better)14.7%#86 of 96, top 90%Vectara Hallucination Leaderboard
GPQA (HELM)79.2%#2 of 57, top 4%HELM Capabilities
LMArena Expert1404LMArena2026-10-08
LMArena Expert1419#108 of 273, top 40%highLMArena2026-10-08

Multimodal

GPT-5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1232#66 of 122, top 55%LMArena2026-10-09
LMArena Vision1208highLMArena2026-10-09
GeoBench81%#4 of 25, top 16%mediumEpoch AI
VPCT66%#4 of 24, top 17%highEpoch AI
VPCT63.2%mediumEpoch AI

Multilingual

GPT-5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1394LMArena2026-10-08
LMArena Non-English1397#110 of 297, top 38%highLMArena2026-10-08
LMArena Chinese1422#125 of 285, top 44%LMArena2026-10-08
LMArena Chinese1422highLMArena2026-10-08
LMArena French1410#112 of 223, top 51%LMArena2026-10-08
LMArena French1408highLMArena2026-10-08
LMArena German1404LMArena2026-10-08
LMArena German1416#82 of 231, top 36%highLMArena2026-10-08
LMArena Japanese1409#51 of 211, top 25%LMArena2026-10-08
LMArena Japanese1402highLMArena2026-10-08
LMArena Korean1360#96 of 213, top 46%LMArena2026-10-08
LMArena Korean1358highLMArena2026-10-08
LMArena Russian1406#99 of 283, top 35%LMArena2026-10-08
LMArena Russian1402highLMArena2026-10-08
LMArena Spanish1399#112 of 226, top 50%LMArena2026-10-08
LMArena Spanish1388highLMArena2026-10-08

Instruction Following

GPT-5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval87.5%#15 of 57, top 27%HELM Capabilities
LMArena Instruction Following1381LMArena2026-10-08
LMArena Instruction Following1388#114 of 298, top 39%highLMArena2026-10-08

Long Context

GPT-5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench97.2%Best of 47mediumEpoch AI
LMArena Longer Query1399#118 of 291, top 41%LMArena2026-10-08
LMArena Longer Query1387highLMArena2026-10-08

Writing & Preference

GPT-5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1403LMArena2026-10-08
LMArena Text1406#114 of 297, top 39%highLMArena2026-10-08
LMArena Creative Writing1365#115 of 295, top 39%LMArena2026-10-08
LMArena Creative Writing1364highLMArena2026-10-08
Short-Story Creative Writing86%Best of 39mediumEpoch AI
EQ-Bench Creative Writing1627#39 of 115, top 34%EQ-Bench
WildBench85.7%#6 of 57, top 11%HELM Capabilities
LMArena Multi-Turn1426#89 of 295, top 31%LMArena2026-10-08
LMArena Multi-Turn1399highLMArena2026-10-08

API pricing by provider

GPT-5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.25$10$0.132026-10-10
openai$1.25$10$0.132026-10-10
openrouter$1.25$10$0.132026-10-10

Compare GPT-5

Other OpenAI models

Frequently asked questions

How good is GPT-5?

GPT-5 by OpenAI ranks 45th of 354 ranked models on the Noometry Index as of October 2026, with a score of 50.9. Its strongest category is long context, where it ranks 2nd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

How much does GPT-5 cost?

GPT-5 costs $1.25 per million input tokens and $10 per million output tokens on OpenAI's own API, with cached input at $0.13.

What is GPT-5's context window?

GPT-5 accepts up to 400K tokens of input and can write up to 128K tokens in one response.

Is GPT-5 open source?

No. GPT-5 is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-5?

GPT-5 generated about 2 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-5's strengths and weaknesses?

Relative to other ranked models, GPT-5 places best in long context, multimodal, math and lowest in instruction following, multilingual, agentic & tool use.

What is GPT-5 best at?

Its best category is long context, where it ranks 2nd on Noometry.