OpenAI, proprietary

GPT-5.5

GPT-5.5 by OpenAI ranks 9th of 354 ranked models on the Noometry Index as of October 2026, with a score of 63.4. Its strongest category is agentic & tool use, where it ranks 6th. API pricing starts at $5 per million input tokens and $30 per million output tokens, with a 1.05M-token context window.

Last verified

Specifications

Noometry rank
#9 of 354
Index score
63.4
Evidence
Confirmed 71 results
Provider
OpenAI
Released
April 23, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1.05M
Max output
128K
Input price
$5 / M
Output price
$30 / M
Blended price
$11.25 / M
Output speed
25 tokens/s Kagi
Value
#204 of 219
Knowledge cutoff
December 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

GPT-5.5 category scores
  1. Coding 58.2
  2. Agentic & Tool Use 50.7
  3. Reasoning 72.8
  4. Math 81.7
  5. Knowledge 64.4
  6. Multimodal 46.9
  7. Multilingual 56.4
  8. Instruction Following 77.5
  9. Long Context 48.3
  10. Writing & Preference 72.7
GPT-5.5 category ranks
CategoryScoreRankResults
Coding58.2#179
Agentic & Tool Use50.7#610
Reasoning72.8#1113
Math81.7#116
Knowledge64.4#174
Multimodal46.9#123
Multilingual56.4#201
Instruction Following77.5#181
Long Context48.3#122
Writing & Preference72.7#135

Strengths and weaknesses

Categories where GPT-5.5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5.5: strongest categories
CategoryScorevs medianRank
Reasoning72.8+49.2#11 of 350, top 4%
Math81.7+45.1#11 of 327, top 4%
Agentic & Tool Use50.7+20.3#6 of 154, top 4%

Weakest categories

GPT-5.5: weakest categories
CategoryScorevs medianRank
Multimodal46.9+8.4#12 of 128, top 10%
Multilingual56.4+9.0#20 of 297, top 7%
Instruction Following77.5+6.3#18 of 305, top 6%

Closest competitors

The models ranked just above and below GPT-5.5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.5
ModelRankScoreBlended $/MSpeed
Claude Fable 5#566.8$2025Compare
GPT-6.1 Sol#665.6$4—Compare
GPT-5.6 Sol#765.0$810Compare
GPT-5.5 Pro#864.3$67.50—Compare
Claude Sonnet 5.5#1061.9$4—Compare
Gemini 3.8 Flash#1161.8$1.50—Compare
GPT-6 Sol#1261.8$4—Compare
Claude Opus 4.8#1360.7$1034Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5.5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified80.6%#2 of 32, top 7%xhighEpoch AI2026-04-24
DeepSWE64.4%highEpoch AI
DeepSWE27%lowEpoch AI
DeepSWE54%mediumEpoch AI
DeepSWE67%#13 of 29, top 45%xhighEpoch AI
FrontierCode43%#15 of 37, top 41%Epoch AI
LMArena WebDev1513#44 of 113, top 39%LMArena2026-10-08
LMArena WebDev1487LMArena2026-10-08
LMArena WebDev1456LMArena2026-10-08
SciCode55.9%highEpoch AI
SciCode51.6%lowEpoch AI
SciCode53.5%mediumEpoch AI
SciCode47.3%noneEpoch AI
SciCode56.1%#23 of 121, top 20%xhighEpoch AI
GSO40.2%#8 of 31, top 26%xhighEpoch AI
WeirdML83.9%highEpoch AI
WeirdML67.2%noneEpoch AI
WeirdML84.9%#6 of 119, top 6%xhighEpoch AI
LMArena Coding1494#27 of 294, top 10%highLMArena2026-10-08
MirrorCode10%#8 of 9, top 89%highEpoch AI2026-08-10
ALE-Bench1,589mediumEpoch AI
ALE-Bench1,128noneEpoch AI
ALE-Bench1,943#9 of 105, top 9%xhighEpoch AI

Agentic & Tool Use

GPT-5.5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench84.7%Best of 41Epoch AI
APEX-Agents55.1%#18 of 49, top 37%Epoch AI
OSWorld 2.013%#5 of 9, top 56%xhighEpoch AI
Remote Labor Index6.3%#5 of 14, top 36%Epoch AI
τ²-bench Banking44.6%#5 of 26, top 20%xhighτ²-bench2026-05-05
DeepResearch Bench54%#4 of 24, top 17%highEpoch AI
DeepResearch Bench48.7%lowEpoch AI
DeepResearch Bench49.6%mediumEpoch AI
PostTrainBench27.2%#8 of 11, top 73%xhighEpoch AI
ExploitBench47.4%#2 of 9, top 23%Epoch AI
GBAEval53.2%#6 of 23, top 27%Epoch AI
GDP.pdf26%#9 of 36, top 25%xhighEpoch AI
LMArena Search1242#3 of 32, top 10%LMArena2026-08-24
Vending-Bench 27,524#13 of 60, top 22%Epoch AI

Reasoning

GPT-5.5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-283.3%highEpoch AI
ARC-AGI-233.3%lowEpoch AI
ARC-AGI-270.4%mediumEpoch AI
ARC-AGI-285%#10 of 83, top 13%xhighEpoch AI
SimpleBench69%#13 of 77, top 17%Epoch AI
Kagi LLM Benchmark88.8%#3 of 99, top 4%Kagi LLM Benchmark
NYT Connections (extended)96.2%#4 of 91, top 5%xhigh reasoningLech Mazur benchmarks
ARC-AGI-194.5%highEpoch AI
ARC-AGI-176.2%lowEpoch AI
ARC-AGI-192.2%mediumEpoch AI
ARC-AGI-195%#15 of 83, top 19%xhighEpoch AI
CritPt25.4%highEpoch AI
CritPt8%lowEpoch AI
CritPt18.6%mediumEpoch AI
CritPt1.4%noneEpoch AI
CritPt27.1%#14 of 134, top 11%xhighEpoch AI
Chess Puzzles26%lowEpoch AI2026-08-07
Chess Puzzles10%noneEpoch AI2026-08-07
Chess Puzzles54%#8 of 129, top 7%xhighEpoch AI2026-04-24
EBR-Bench34.3%#9 of 24, top 38%xhighEpoch AI2026-07-27
LMArena Hard Prompts1489#17 of 297, top 6%highLMArena2026-10-08
Mystery Game Puzzles52%highEpoch AI2026-07-27
Mystery Game Puzzles28%lowEpoch AI2026-08-28
Mystery Game Puzzles18%noneEpoch AI2026-08-27
Mystery Game Puzzles56%#8 of 74, top 11%xhighEpoch AI2026-07-24
DTBench96%#12 of 151, top 8%xhighEpoch AI
LMCA54.3%#12 of 125, top 10%xhighEpoch AI
Surface Evolver Bench88.1%#4 of 25, top 16%highEpoch AI
Surface Evolver Bench81.3%mediumEpoch AI
Bench to the Future 30.14#3 of 10, top 30%highEpoch AI
Epoch Capabilities Index159.1#12 of 213, top 6%Epoch AI2026-04-23
ForecastBench60.6#25 of 72, top 35%Epoch AI

Math

GPT-5.5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)85.3%#12 of 81, top 15%xhighEpoch AI2026-06-11
FrontierMath Tier 472.5%#12 of 63, top 20%xhighEpoch AI2026-06-11
MathArena Final-Answer Competitions94.3%Best of 29xhighMathArena
OTIS Mock AIME 2024-202584.4%lowEpoch AI2026-08-07
OTIS Mock AIME 2024-202557.8%noneEpoch AI2026-08-07
OTIS Mock AIME 2024-2025100%#5 of 173, top 3%xhighEpoch AI2026-04-24
ProofBench50%#30 of 77, top 39%xhighEpoch AI
LMArena Math1486#23 of 285, top 9%LMArena2026-10-08
FrontierMath (Feb 2025 set)51.7%#2 of 68, top 3%xhighEpoch AI2026-04-23
FrontierMath Erdős0%#6 of 7, top 86%xhighEpoch AI2026-08-28
FrontierMath Tier 4 (v1)35.4%#4 of 55, top 8%xhighEpoch AI2026-04-23

Knowledge

GPT-5.5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond90.7%lowEpoch AI2026-05-05
GPQA Diamond77.3%noneEpoch AI2026-08-07
GPQA Diamond94%#10 of 186, top 6%xhighEpoch AI2026-04-24
SimpleQA Verified63%#13 of 77, top 17%xhighEpoch AI2026-08-27
Vectara Hallucination Rate (lower is better)9.3%#46 of 96, top 48%Vectara Hallucination Leaderboard
LMArena Expert1508#17 of 273, top 7%highLMArena2026-10-08

Multimodal

GPT-5.5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1297#17 of 122, top 14%LMArena2026-10-09
Blueprint-Bench 236.2%#8 of 31, top 26%Epoch AI
Furniture Assembly44.2%#11 of 31, top 36%xhighEpoch AI2026-09-10
LMArena Document1486#6 of 38, top 16%LMArena2026-09-13

Multilingual

GPT-5.5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1467#20 of 297, top 7%highLMArena2026-10-08
LMArena Chinese1533#11 of 285, top 4%LMArena2026-10-08
LMArena French1486#24 of 223, top 11%LMArena2026-10-08
LMArena German1480#20 of 231, top 9%LMArena2026-10-08
LMArena Japanese1498#7 of 211, top 4%highLMArena2026-10-08
LMArena Korean1460#10 of 213, top 5%highLMArena2026-10-08
LMArena Russian1473#24 of 283, top 9%LMArena2026-10-08
LMArena Spanish1468#28 of 226, top 13%highLMArena2026-10-08

Instruction Following

GPT-5.5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1479#15 of 298, top 6%highLMArena2026-10-08

Long Context

GPT-5.5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench Life22.2%Best of 13highEpoch AI
LMArena Longer Query1484#15 of 291, top 6%highLMArena2026-10-08

Writing & Preference

GPT-5.5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1472#22 of 297, top 8%highLMArena2026-10-08
LMArena Creative Writing1455#22 of 295, top 8%highLMArena2026-10-08
EQ-Bench Creative Writing1844#16 of 115, top 14%EQ-Bench
EQ-Bench 41315#4 of 28, top 15%EQ-Bench
LMArena Multi-Turn1476#24 of 295, top 9%highLMArena2026-10-08

API pricing by provider

GPT-5.5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$5$30$0.502026-10-10
bedrock$5.50$33$0.552026-10-10
openai$5$30$0.502026-10-10
openrouter$5$30$0.502026-10-10

Compare GPT-5.5

Other OpenAI models

Frequently asked questions

How good is GPT-5.5?

GPT-5.5 by OpenAI ranks 9th of 354 ranked models on the Noometry Index as of October 2026, with a score of 63.4. Its strongest category is agentic & tool use, where it ranks 6th. API pricing starts at $5 per million input tokens and $30 per million output tokens, with a 1.05M-token context window.

How much does GPT-5.5 cost?

GPT-5.5 costs $5 per million input tokens and $30 per million output tokens on OpenAI's own API, with cached input at $0.50.

What is GPT-5.5's context window?

GPT-5.5 accepts up to 1.05M tokens of input and can write up to 128K tokens in one response.

Is GPT-5.5 open source?

No. GPT-5.5 is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-5.5?

GPT-5.5 generated about 25 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-5.5's strengths and weaknesses?

Relative to other ranked models, GPT-5.5 places best in reasoning, math, agentic & tool use and lowest in multimodal, multilingual, instruction following.

What is GPT-5.5 best at?

Its best category is agentic & tool use, where it ranks 6th on Noometry.