OpenAI, proprietary
GPT-5.5
GPT-5.5 by OpenAI ranks 9th of 354 ranked models on the Noometry Index as of October 2026, with a score of 63.4. Its strongest category is agentic & tool use, where it ranks 6th. API pricing starts at $5 per million input tokens and $30 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #9 of 354
- Index score
- 63.4
- Evidence
- Confirmed 71 results
- Provider
- OpenAI
- Released
- April 23, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1.05M
- Max output
- 128K
- Input price
- $5 / M
- Output price
- $30 / M
- Blended price
- $11.25 / M
- Output speed
- 25 tokens/s Kagi
- Value
- #204 of 219
- Knowledge cutoff
- December 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 58.2
- Agentic & Tool Use 50.7
- Reasoning 72.8
- Math 81.7
- Knowledge 64.4
- Multimodal 46.9
- Multilingual 56.4
- Instruction Following 77.5
- Long Context 48.3
- Writing & Preference 72.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 58.2 | #17 | 9 |
| Agentic & Tool Use | 50.7 | #6 | 10 |
| Reasoning | 72.8 | #11 | 13 |
| Math | 81.7 | #11 | 6 |
| Knowledge | 64.4 | #17 | 4 |
| Multimodal | 46.9 | #12 | 3 |
| Multilingual | 56.4 | #20 | 1 |
| Instruction Following | 77.5 | #18 | 1 |
| Long Context | 48.3 | #12 | 2 |
| Writing & Preference | 72.7 | #13 | 5 |
Strengths and weaknesses
Categories where GPT-5.5 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 72.8 | +49.2 | #11 of 350, top 4% |
| Math | 81.7 | +45.1 | #11 of 327, top 4% |
| Agentic & Tool Use | 50.7 | +20.3 | #6 of 154, top 4% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 46.9 | +8.4 | #12 of 128, top 10% |
| Multilingual | 56.4 | +9.0 | #20 of 297, top 7% |
| Instruction Following | 77.5 | +6.3 | #18 of 305, top 6% |
Closest competitors
The models ranked just above and below GPT-5.5. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude Fable 5 | #5 | 66.8 | $20 | 25 | Compare |
| GPT-6.1 Sol | #6 | 65.6 | $4 | — | Compare |
| GPT-5.6 Sol | #7 | 65.0 | $8 | 10 | Compare |
| GPT-5.5 Pro | #8 | 64.3 | $67.50 | — | Compare |
| Claude Sonnet 5.5 | #10 | 61.9 | $4 | — | Compare |
| Gemini 3.8 Flash | #11 | 61.8 | $1.50 | — | Compare |
| GPT-6 Sol | #12 | 61.8 | $4 | — | Compare |
| Claude Opus 4.8 | #13 | 60.7 | $10 | 34 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 80.6% | #2 of 32, top 7% | xhigh | Epoch AI | 2026-04-24 |
| DeepSWE | 64.4% | high | Epoch AI | ||
| DeepSWE | 27% | low | Epoch AI | ||
| DeepSWE | 54% | medium | Epoch AI | ||
| DeepSWE | 67% | #13 of 29, top 45% | xhigh | Epoch AI | |
| FrontierCode | 43% | #15 of 37, top 41% | Epoch AI | ||
| LMArena WebDev | 1513 | #44 of 113, top 39% | LMArena | 2026-10-08 | |
| LMArena WebDev | 1487 | LMArena | 2026-10-08 | ||
| LMArena WebDev | 1456 | LMArena | 2026-10-08 | ||
| SciCode | 55.9% | high | Epoch AI | ||
| SciCode | 51.6% | low | Epoch AI | ||
| SciCode | 53.5% | medium | Epoch AI | ||
| SciCode | 47.3% | none | Epoch AI | ||
| SciCode | 56.1% | #23 of 121, top 20% | xhigh | Epoch AI | |
| GSO | 40.2% | #8 of 31, top 26% | xhigh | Epoch AI | |
| WeirdML | 83.9% | high | Epoch AI | ||
| WeirdML | 67.2% | none | Epoch AI | ||
| WeirdML | 84.9% | #6 of 119, top 6% | xhigh | Epoch AI | |
| LMArena Coding | 1494 | #27 of 294, top 10% | high | LMArena | 2026-10-08 |
| MirrorCode | 10% | #8 of 9, top 89% | high | Epoch AI | 2026-08-10 |
| ALE-Bench | 1,589 | medium | Epoch AI | ||
| ALE-Bench | 1,128 | none | Epoch AI | ||
| ALE-Bench | 1,943 | #9 of 105, top 9% | xhigh | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 84.7% | Best of 41 | Epoch AI | ||
| APEX-Agents | 55.1% | #18 of 49, top 37% | Epoch AI | ||
| OSWorld 2.0 | 13% | #5 of 9, top 56% | xhigh | Epoch AI | |
| Remote Labor Index | 6.3% | #5 of 14, top 36% | Epoch AI | ||
| τ²-bench Banking | 44.6% | #5 of 26, top 20% | xhigh | τ²-bench | 2026-05-05 |
| DeepResearch Bench | 54% | #4 of 24, top 17% | high | Epoch AI | |
| DeepResearch Bench | 48.7% | low | Epoch AI | ||
| DeepResearch Bench | 49.6% | medium | Epoch AI | ||
| PostTrainBench | 27.2% | #8 of 11, top 73% | xhigh | Epoch AI | |
| ExploitBench | 47.4% | #2 of 9, top 23% | Epoch AI | ||
| GBAEval | 53.2% | #6 of 23, top 27% | Epoch AI | ||
| GDP.pdf | 26% | #9 of 36, top 25% | xhigh | Epoch AI | |
| LMArena Search | 1242 | #3 of 32, top 10% | LMArena | 2026-08-24 | |
| Vending-Bench 2 | 7,524 | #13 of 60, top 22% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 83.3% | high | Epoch AI | ||
| ARC-AGI-2 | 33.3% | low | Epoch AI | ||
| ARC-AGI-2 | 70.4% | medium | Epoch AI | ||
| ARC-AGI-2 | 85% | #10 of 83, top 13% | xhigh | Epoch AI | |
| SimpleBench | 69% | #13 of 77, top 17% | Epoch AI | ||
| Kagi LLM Benchmark | 88.8% | #3 of 99, top 4% | Kagi LLM Benchmark | ||
| NYT Connections (extended) | 96.2% | #4 of 91, top 5% | xhigh reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 94.5% | high | Epoch AI | ||
| ARC-AGI-1 | 76.2% | low | Epoch AI | ||
| ARC-AGI-1 | 92.2% | medium | Epoch AI | ||
| ARC-AGI-1 | 95% | #15 of 83, top 19% | xhigh | Epoch AI | |
| CritPt | 25.4% | high | Epoch AI | ||
| CritPt | 8% | low | Epoch AI | ||
| CritPt | 18.6% | medium | Epoch AI | ||
| CritPt | 1.4% | none | Epoch AI | ||
| CritPt | 27.1% | #14 of 134, top 11% | xhigh | Epoch AI | |
| Chess Puzzles | 26% | low | Epoch AI | 2026-08-07 | |
| Chess Puzzles | 10% | none | Epoch AI | 2026-08-07 | |
| Chess Puzzles | 54% | #8 of 129, top 7% | xhigh | Epoch AI | 2026-04-24 |
| EBR-Bench | 34.3% | #9 of 24, top 38% | xhigh | Epoch AI | 2026-07-27 |
| LMArena Hard Prompts | 1489 | #17 of 297, top 6% | high | LMArena | 2026-10-08 |
| Mystery Game Puzzles | 52% | high | Epoch AI | 2026-07-27 | |
| Mystery Game Puzzles | 28% | low | Epoch AI | 2026-08-28 | |
| Mystery Game Puzzles | 18% | none | Epoch AI | 2026-08-27 | |
| Mystery Game Puzzles | 56% | #8 of 74, top 11% | xhigh | Epoch AI | 2026-07-24 |
| DTBench | 96% | #12 of 151, top 8% | xhigh | Epoch AI | |
| LMCA | 54.3% | #12 of 125, top 10% | xhigh | Epoch AI | |
| Surface Evolver Bench | 88.1% | #4 of 25, top 16% | high | Epoch AI | |
| Surface Evolver Bench | 81.3% | medium | Epoch AI | ||
| Bench to the Future 3 | 0.14 | #3 of 10, top 30% | high | Epoch AI | |
| Epoch Capabilities Index | 159.1 | #12 of 213, top 6% | Epoch AI | 2026-04-23 | |
| ForecastBench | 60.6 | #25 of 72, top 35% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 85.3% | #12 of 81, top 15% | xhigh | Epoch AI | 2026-06-11 |
| FrontierMath Tier 4 | 72.5% | #12 of 63, top 20% | xhigh | Epoch AI | 2026-06-11 |
| MathArena Final-Answer Competitions | 94.3% | Best of 29 | xhigh | MathArena | |
| OTIS Mock AIME 2024-2025 | 84.4% | low | Epoch AI | 2026-08-07 | |
| OTIS Mock AIME 2024-2025 | 57.8% | none | Epoch AI | 2026-08-07 | |
| OTIS Mock AIME 2024-2025 | 100% | #5 of 173, top 3% | xhigh | Epoch AI | 2026-04-24 |
| ProofBench | 50% | #30 of 77, top 39% | xhigh | Epoch AI | |
| LMArena Math | 1486 | #23 of 285, top 9% | LMArena | 2026-10-08 | |
| FrontierMath (Feb 2025 set) | 51.7% | #2 of 68, top 3% | xhigh | Epoch AI | 2026-04-23 |
| FrontierMath Erdős | 0% | #6 of 7, top 86% | xhigh | Epoch AI | 2026-08-28 |
| FrontierMath Tier 4 (v1) | 35.4% | #4 of 55, top 8% | xhigh | Epoch AI | 2026-04-23 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 90.7% | low | Epoch AI | 2026-05-05 | |
| GPQA Diamond | 77.3% | none | Epoch AI | 2026-08-07 | |
| GPQA Diamond | 94% | #10 of 186, top 6% | xhigh | Epoch AI | 2026-04-24 |
| SimpleQA Verified | 63% | #13 of 77, top 17% | xhigh | Epoch AI | 2026-08-27 |
| Vectara Hallucination Rate (lower is better) | 9.3% | #46 of 96, top 48% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1508 | #17 of 273, top 7% | high | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1297 | #17 of 122, top 14% | LMArena | 2026-10-09 | |
| Blueprint-Bench 2 | 36.2% | #8 of 31, top 26% | Epoch AI | ||
| Furniture Assembly | 44.2% | #11 of 31, top 36% | xhigh | Epoch AI | 2026-09-10 |
| LMArena Document | 1486 | #6 of 38, top 16% | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1467 | #20 of 297, top 7% | high | LMArena | 2026-10-08 |
| LMArena Chinese | 1533 | #11 of 285, top 4% | LMArena | 2026-10-08 | |
| LMArena French | 1486 | #24 of 223, top 11% | LMArena | 2026-10-08 | |
| LMArena German | 1480 | #20 of 231, top 9% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1498 | #7 of 211, top 4% | high | LMArena | 2026-10-08 |
| LMArena Korean | 1460 | #10 of 213, top 5% | high | LMArena | 2026-10-08 |
| LMArena Russian | 1473 | #24 of 283, top 9% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1468 | #28 of 226, top 13% | high | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1479 | #15 of 298, top 6% | high | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CL-bench Life | 22.2% | Best of 13 | high | Epoch AI | |
| LMArena Longer Query | 1484 | #15 of 291, top 6% | high | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1472 | #22 of 297, top 8% | high | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1455 | #22 of 295, top 8% | high | LMArena | 2026-10-08 |
| EQ-Bench Creative Writing | 1844 | #16 of 115, top 14% | EQ-Bench | ||
| EQ-Bench 4 | 1315 | #4 of 28, top 15% | EQ-Bench | ||
| LMArena Multi-Turn | 1476 | #24 of 295, top 9% | high | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $5 | $30 | $0.50 | 2026-10-10 |
| bedrock | $5.50 | $33 | $0.55 | 2026-10-10 |
| openai | $5 | $30 | $0.50 | 2026-10-10 |
| openrouter | $5 | $30 | $0.50 | 2026-10-10 |
Compare GPT-5.5
- GPT-5.5 vs GPT-5.4
- GPT-5.5 vs GPT-5.5 Pro
- GPT-5.5 vs Claude Sonnet 5.5
- GPT-5.5 vs GPT-5.6 Sol
- GPT-5.5 vs Gemini 3.8 Flash
- GPT-5.5 vs GPT-6.1 Sol
- GPT-5.5 vs GPT-6 Sol
- GPT-5.5 vs Claude Fable 5.1
- GPT-5.5 vs Kimi K3
- GPT-5.5 vs Grok 4.6
- GPT-5.5 vs Qwen3.8 Max
- GPT-5.5 vs GLM-5.3
- GPT-5.5 vs Muse Spark 1.3
- GPT-5.5 vs DeepSeek V4 Pro
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
- GPT-5.4 Pro58.9
Frequently asked questions
How good is GPT-5.5?
GPT-5.5 by OpenAI ranks 9th of 354 ranked models on the Noometry Index as of October 2026, with a score of 63.4. Its strongest category is agentic & tool use, where it ranks 6th. API pricing starts at $5 per million input tokens and $30 per million output tokens, with a 1.05M-token context window.
How much does GPT-5.5 cost?
GPT-5.5 costs $5 per million input tokens and $30 per million output tokens on OpenAI's own API, with cached input at $0.50.
What is GPT-5.5's context window?
GPT-5.5 accepts up to 1.05M tokens of input and can write up to 128K tokens in one response.
Is GPT-5.5 open source?
No. GPT-5.5 is proprietary and available only through OpenAI's API and partner platforms.
How fast is GPT-5.5?
GPT-5.5 generated about 25 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are GPT-5.5's strengths and weaknesses?
Relative to other ranked models, GPT-5.5 places best in reasoning, math, agentic & tool use and lowest in multimodal, multilingual, instruction following.
What is GPT-5.5 best at?
Its best category is agentic & tool use, where it ranks 6th on Noometry.