OpenAI, proprietary

GPT-5.6 Sol

GPT-5.6 Sol by OpenAI ranks 7th of 354 ranked models on the Noometry Index as of October 2026, with a score of 65.0. Its strongest category is agentic & tool use, where it ranks 7th. API pricing starts at $4 per million input tokens and $20 per million output tokens, with a 1.05M-token context window.

Last verified

Specifications

Noometry rank
#7 of 354
Index score
65.0
Evidence
Confirmed 65 results
Provider
OpenAI
Released
July 9, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1.05M
Max output
128K
Input price
$4 / M
Output price
$20 / M
Blended price
$8 / M
Output speed
10 tokens/s Kagi
Value
#193 of 219
Knowledge cutoff
February 2026
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

GPT-5.6 Sol category scores
  1. Coding 65.1
  2. Agentic & Tool Use 50.3
  3. Reasoning 74.8
  4. Math 85.6
  5. Knowledge 64.3
  6. Multimodal 48.6
  7. Multilingual 55.3
  8. Instruction Following 77.7
  9. Long Context 45.4
  10. Writing & Preference 73.3
GPT-5.6 Sol category ranks
CategoryScoreRankResults
Coding65.1#710
Agentic & Tool Use50.3#77
Reasoning74.8#814
Math85.6#95
Knowledge64.3#184
Multimodal48.6#93
Multilingual55.3#321
Instruction Following77.7#161
Long Context45.4#421
Writing & Preference73.3#125

Strengths and weaknesses

Categories where GPT-5.6 Sol places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5.6 Sol: strongest categories
CategoryScorevs medianRank
Coding65.1+26.3#7 of 340, top 3%
Reasoning74.8+51.2#8 of 350, top 3%
Math85.6+49.0#9 of 327, top 3%

Weakest categories

GPT-5.6 Sol: weakest categories
CategoryScorevs medianRank
Long Context45.4+4.4#42 of 296, top 15%
Multilingual55.3+7.9#32 of 297, top 11%
Multimodal48.6+10.1#9 of 128, top 8%

Closest competitors

The models ranked just above and below GPT-5.6 Sol. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.6 Sol
ModelRankScoreBlended $/MSpeed
Claude Opus 5.5#368.6$8—Compare
Claude Opus 5#467.8$10—Compare
Claude Fable 5#566.8$2025Compare
GPT-6.1 Sol#665.6$4—Compare
GPT-5.5 Pro#864.3$67.50—Compare
GPT-5.5#963.4$11.2525Compare
Claude Sonnet 5.5#1061.9$4—Compare
Gemini 3.8 Flash#1161.8$1.50—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5.6 Sol Coding benchmark results
BenchmarkScorePositionSettingSourceDate
DeepSWE69.4%highEpoch AI
DeepSWE45.4%lowEpoch AI
DeepSWE72.7%#5 of 29, top 18%maxEpoch AI
DeepSWE61.1%mediumEpoch AI
DeepSWE70.7%xhighEpoch AI
FrontierCode47.5%#11 of 37, top 30%Epoch AI
CursorBench35.7%highEpoch AI
CursorBench24.6%lowEpoch AI
CursorBench41.7%#7 of 14, top 50%maxEpoch AI
CursorBench31.1%mediumEpoch AI
CursorBench37.7%xhighEpoch AI
LMArena WebDev1618#19 of 113, top 17%LMArena2026-10-08
FrontierSWE32.2%#8 of 18, top 45%maxEpoch AI
SciCode56.9%highEpoch AI
SciCode55.4%lowEpoch AI
SciCode57.1%#16 of 121, top 14%maxEpoch AI
SciCode56.5%mediumEpoch AI
SciCode47.1%noneEpoch AI
SciCode56%xhighEpoch AI
GSO76.5%#4 of 31, top 13%Epoch AI
WeirdML88.8%highEpoch AI
WeirdML87%maxEpoch AI
WeirdML89.4%#5 of 119, top 5%promaxEpoch AI
LMArena Coding1498#21 of 294, top 8%xhighLMArena2026-10-08
MirrorCode20%#6 of 9, top 67%highEpoch AI2026-08-10
ALE-Bench2,177#3 of 105, top 3%maxEpoch AI

Agentic & Tool Use

GPT-5.6 Sol Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents51.4%#24 of 49, top 49%promaxEpoch AI
OSWorld 2.027.3%#2 of 9, top 23%maxEpoch AI
τ²-bench Banking46.9%#4 of 26, top 16%xhighτ²-bench2026-08-04
PostTrainBench36.2%#2 of 11, top 19%maxEpoch AI
BALROG60%#3 of 35, top 9%maxEpoch AI
GBAEval52.6%#7 of 23, top 31%Epoch AI
GDP.pdf30.7%#3 of 36, top 9%Epoch AI
GDP.pdf30.7%maxEpoch AI
LMArena Search1257Best of 32xhighLMArena2026-08-24
Vending-Bench 29,619#7 of 60, top 12%Epoch AI

Reasoning

GPT-5.6 Sol Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-285.4%highEpoch AI
ARC-AGI-242.5%lowEpoch AI
ARC-AGI-292.5%#4 of 83, top 5%maxEpoch AI
ARC-AGI-267.1%mediumEpoch AI
ARC-AGI-290%xhighEpoch AI
SimpleBench64.8%Epoch AI
SimpleBench71.7%#10 of 77, top 13%prounknownEpoch AI
SimpleBench71.7%proxhighEpoch AI
SimpleBench64.8%xhighEpoch AI
Kagi LLM Benchmark67%#30 of 99, top 31%Kagi LLM Benchmark
NYT Connections (extended)93.8%#9 of 91, top 10%xhigh reasoningLech Mazur benchmarks
ARC-AGI-197%highEpoch AI
ARC-AGI-174.5%lowEpoch AI
ARC-AGI-196.5%maxEpoch AI
ARC-AGI-192.5%mediumEpoch AI
ARC-AGI-197.5%#9 of 83, top 11%xhighEpoch AI
CritPt25.7%highEpoch AI
CritPt14.9%lowEpoch AI
CritPt32.3%Best of 134maxEpoch AI
CritPt22.9%mediumEpoch AI
CritPt5.1%noneEpoch AI
CritPt28.6%xhighEpoch AI
Chess Puzzles27%lowEpoch AI2026-08-07
Chess Puzzles55%maxEpoch AI2026-07-09
Chess Puzzles7%noneEpoch AI2026-08-07
Chess Puzzles64%#3 of 129, top 3%promaxEpoch AI2026-07-10
EnigmaEval37.1%#2 of 38, top 6%highEpoch AI
EBR-Bench44.8%#7 of 24, top 30%maxEpoch AI2026-09-04
LMArena Hard Prompts1484#26 of 297, top 9%xhighLMArena2026-10-08
Mystery Game Puzzles26%lowEpoch AI2026-08-27
Mystery Game Puzzles58%#7 of 74, top 10%maxEpoch AI2026-07-28
Mystery Game Puzzles33%noneEpoch AI2026-08-27
DTBench95.5%highEpoch AI
DTBench93.9%lowEpoch AI
DTBench95.5%maxEpoch AI
DTBench95.2%mediumEpoch AI
DTBench83.7%noneEpoch AI
DTBench96%#14 of 151, top 10%promaxEpoch AI
DTBench95.7%xhighEpoch AI
LMCA56.9%highEpoch AI
LMCA55.9%lowEpoch AI
LMCA58.4%maxEpoch AI
LMCA56.4%mediumEpoch AI
LMCA52.8%noneEpoch AI
LMCA59.2%#6 of 125, top 5%promaxEpoch AI
LMCA58.5%xhighEpoch AI
Surface Evolver Bench93.1%#3 of 25, top 12%xhighEpoch AI
Bench to the Future 30.14highEpoch AI
Bench to the Future 30.14#7 of 10, top 70%maxEpoch AI
Bench to the Future 30.14mediumEpoch AI
Epoch Capabilities Index161.66#10 of 213, top 5%Epoch AI2026-07-09

Math

GPT-5.6 Sol Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)89.1%#6 of 81, top 8%maxEpoch AI2026-07-09
FrontierMath Tier 482.9%#7 of 63, top 12%maxEpoch AI2026-07-09
FrontierMath Tier 480.5%promaxEpoch AI2026-07-09
OTIS Mock AIME 2024-202595.6%lowEpoch AI2026-08-07
OTIS Mock AIME 2024-2025100%#7 of 173, top 5%maxEpoch AI2026-07-09
OTIS Mock AIME 2024-202568.9%noneEpoch AI2026-08-07
ProofBench83%#10 of 77, top 13%maxEpoch AI
LMArena Math1474#35 of 285, top 13%xhighLMArena2026-10-08
FrontierMath Erdős0%#7 of 7, top 100%maxEpoch AI2026-08-28

Knowledge

GPT-5.6 Sol Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond89.9%lowEpoch AI2026-08-07
GPQA Diamond93.5%#14 of 186, top 8%maxEpoch AI2026-07-09
GPQA Diamond82.8%noneEpoch AI2026-08-07
SimpleQA Verified69.7%#8 of 77, top 11%maxEpoch AI2026-08-10
Vectara Hallucination Rate (lower is better)12.4%#78 of 96, top 82%Vectara Hallucination Leaderboard
LMArena Expert1516#12 of 273, top 5%xhighLMArena2026-10-08

Multimodal

GPT-5.6 Sol Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1281#28 of 122, top 23%xhighLMArena2026-10-09
Blueprint-Bench 233.6%#10 of 31, top 33%Epoch AI
Furniture Assembly56.7%#8 of 31, top 26%maxEpoch AI2026-09-10
LMArena Document1483#7 of 38, top 19%xhighLMArena2026-09-13

Multilingual

GPT-5.6 Sol Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1452#31 of 297, top 11%xhighLMArena2026-10-08
LMArena Chinese1527#21 of 285, top 8%xhighLMArena2026-10-08
LMArena French1477#29 of 223, top 14%xhighLMArena2026-10-08
LMArena German1476#22 of 231, top 10%xhighLMArena2026-10-08
LMArena Japanese1471#18 of 211, top 9%xhighLMArena2026-10-08
LMArena Korean1442#24 of 213, top 12%xhighLMArena2026-10-08
LMArena Russian1468#28 of 283, top 10%xhighLMArena2026-10-08
LMArena Spanish1441#64 of 226, top 29%xhighLMArena2026-10-08

Instruction Following

GPT-5.6 Sol Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1482#13 of 298, top 5%xhighLMArena2026-10-08

Long Context

GPT-5.6 Sol Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1480#23 of 291, top 8%xhighLMArena2026-10-08

Writing & Preference

GPT-5.6 Sol Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1457#37 of 297, top 13%xhighLMArena2026-10-08
LMArena Creative Writing1448#30 of 295, top 11%xhighLMArena2026-10-08
EQ-Bench Creative Writing1972#9 of 115, top 8%EQ-Bench
EQ-Bench 41250#9 of 28, top 33%EQ-Bench
LMArena Multi-Turn1460#38 of 295, top 13%xhighLMArena2026-10-08

API pricing by provider

GPT-5.6 Sol API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$4$20$0.502026-10-10
bedrock$4$20$0.402026-10-10
openai$4$20$0.402026-10-10
openrouter$2$10$0.202026-10-10

Compare GPT-5.6 Sol

Other OpenAI models

Frequently asked questions

How good is GPT-5.6 Sol?

GPT-5.6 Sol by OpenAI ranks 7th of 354 ranked models on the Noometry Index as of October 2026, with a score of 65.0. Its strongest category is agentic & tool use, where it ranks 7th. API pricing starts at $4 per million input tokens and $20 per million output tokens, with a 1.05M-token context window.

How much does GPT-5.6 Sol cost?

GPT-5.6 Sol costs $4 per million input tokens and $20 per million output tokens on OpenAI's own API, with cached input at $0.40.

What is GPT-5.6 Sol's context window?

GPT-5.6 Sol accepts up to 1.05M tokens of input and can write up to 128K tokens in one response.

Is GPT-5.6 Sol open source?

No. GPT-5.6 Sol is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-5.6 Sol?

GPT-5.6 Sol generated about 10 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-5.6 Sol's strengths and weaknesses?

Relative to other ranked models, GPT-5.6 Sol places best in coding, reasoning, math and lowest in long context, multilingual, multimodal.

What is GPT-5.6 Sol best at?

Its best category is agentic & tool use, where it ranks 7th on Noometry.