OpenAI, proprietary

GPT-5.1

GPT-5.1 by OpenAI ranks 53rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 49.0. Its strongest category is instruction following, where it ranks 1st. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

Last verified

Specifications

Noometry rank
#53 of 354
Index score
49.0
Evidence
Confirmed 63 results
Provider
OpenAI
Released
November 13, 2025
Weights
Proprietary
Reasoning
Yes
Context window
400K
Max output
128K
Input price
$1.25 / M
Output price
$10 / M
Blended price
$3.44 / M
Output speed
Not measured
Value
#165 of 219
Knowledge cutoff
September 2024
Input
text, image

Category scores

Each category score combines every public result we have in that category.

GPT-5.1 category scores
  1. Coding 46.4
  2. Agentic & Tool Use 32.7
  3. Reasoning 39.8
  4. Math 52.2
  5. Knowledge 50.6
  6. Multimodal 44.8
  7. Multilingual 53.8
  8. Instruction Following 83.9
  9. Long Context 47.6
  10. Writing & Preference 64.5
GPT-5.1 category ranks
CategoryScoreRankResults
Coding46.4#668
Agentic & Tool Use32.7#602
Reasoning39.8#5812
Math52.2#514
Knowledge50.6#717
Multimodal44.8#192
Multilingual53.8#561
Instruction Following83.9#13
Long Context47.6#143
Writing & Preference64.5#555

Strengths and weaknesses

Categories where GPT-5.1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5.1: strongest categories
CategoryScorevs medianRank
Instruction Following83.9+12.6#1 of 305, top 1%
Long Context47.6+6.7#14 of 296, top 5%
Multimodal44.8+6.2#19 of 128, top 15%

Weakest categories

GPT-5.1: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use32.7+2.4#60 of 154, top 39%
Knowledge50.6+13.3#71 of 314, top 23%
Coding46.4+7.7#66 of 340, top 20%

Closest competitors

The models ranked just above and below GPT-5.1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.1
ModelRankScoreBlended $/MSpeed
MiMo-V2.6-Pro#4950.3$0.54—Compare
Claude Sonnet 4.6#5050.3$6—Compare
Muse Spark 1.1#5149.9$2—Compare
Claude Haiku 5.5#5249.5$0.20—Compare
Grok 4.20 (Non-Reasoning)#5448.6$1.5661Compare
MiMo-V2.6-Flash#5548.5$0.18—Compare
Grok 4#5648.1—1Compare
Kimi K2.5#5748.1$0.9066Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5.1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified68%#25 of 32, top 79%highEpoch AI2026-02-18
SWE-bench Verified (bash only)66%#14 of 39, top 36%mediumSWE-bench2025-11-20
LMArena WebDev1395#75 of 113, top 67%mediumLMArena2026-10-08
SciCode43.3%#65 of 121, top 54%Epoch AI
GSO13.7%#16 of 31, top 52%Epoch AI
GSO13.7%highEpoch AI
WeirdML60.8%#29 of 119, top 25%highEpoch AI
LiveBench Coding72.5%#5 of 39, top 13%highEpoch AI
LMArena Coding1454#81 of 294, top 28%highLMArena2026-10-08
ALE-Bench1,192#30 of 105, top 29%highEpoch AI

Agentic & Tool Use

GPT-5.1 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench47.6%Epoch AI
Terminal-Bench47.6%#18 of 41, top 44%mediumEpoch AI
DeepResearch Bench42.8%#19 of 24, top 80%lowEpoch AI
LMArena Search1199#14 of 32, top 44%LMArena2026-08-24
Vending-Bench 21,473#44 of 60, top 74%Epoch AI

Reasoning

GPT-5.1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-217.6%#44 of 83, top 54%highEpoch AI
ARC-AGI-21.9%lowEpoch AI
ARC-AGI-26.5%mediumEpoch AI
ARC-AGI-20.4%noneEpoch AI
SimpleBench53.2%#37 of 77, top 49%highEpoch AI
ARC-AGI-172.8%#42 of 83, top 51%highEpoch AI
ARC-AGI-133.2%lowEpoch AI
ARC-AGI-157.7%mediumEpoch AI
ARC-AGI-15.8%noneEpoch AI
CritPt4.9%#58 of 134, top 44%Epoch AI
Chess Puzzles32%#32 of 129, top 25%highEpoch AI2025-12-08
Chess Puzzles14%lowEpoch AI2026-08-07
Chess Puzzles17%noneEpoch AI2026-08-07
EnigmaEval11.2%#12 of 38, top 32%Epoch AI
EnigmaEval1.9%noneEpoch AI
LiveBench Reasoning95.8%Best of 39highEpoch AI
LMArena Hard Prompts1457#52 of 297, top 18%highLMArena2026-10-08
Mystery Game Puzzles19%#45 of 74, top 61%lowEpoch AI2026-08-27
Mystery Game Puzzles16%mediumEpoch AI2026-08-27
Mystery Game Puzzles15%noneEpoch AI2026-08-27
DTBench90.1%#37 of 151, top 25%highEpoch AI
LiveBench Data Analysis72.1%#3 of 39, top 8%highEpoch AI
LMCA43.9%#39 of 125, top 32%highEpoch AI
Epoch Capabilities Index149.64#56 of 213, top 27%Epoch AI2025-11-13
ForecastBench58.1#51 of 72, top 71%Epoch AI
LiveBench78.8%#2 of 39, top 6%highEpoch AI

Math

GPT-5.1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202588.6%#58 of 173, top 34%highEpoch AI2025-11-13
OTIS Mock AIME 2024-202563.9%lowEpoch AI2025-11-25
OTIS Mock AIME 2024-202585.6%mediumEpoch AI2025-11-17
OTIS Mock AIME 2024-202537.8%noneEpoch AI2026-08-07
Omni-MATH46.4%#21 of 57, top 37%HELM Capabilities
LiveBench Math94.5%Best of 39highEpoch AI
LMArena Math1447#65 of 285, top 23%highLMArena2026-10-08
FrontierMath (Feb 2025 set)31%#18 of 68, top 27%highEpoch AI2025-11-13
FrontierMath (Feb 2025 set)17.3%lowEpoch AI2025-11-25
FrontierMath (Feb 2025 set)26.9%mediumEpoch AI2025-11-17
FrontierMath (Feb 2025 set)2.1%noneEpoch AI2025-11-25
FrontierMath Tier 4 (v1)12.5%#19 of 55, top 35%highEpoch AI2025-11-13
FrontierMath Tier 4 (v1)4.2%mediumEpoch AI2025-11-17

Knowledge

GPT-5.1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond87.6%#52 of 186, top 28%highEpoch AI2025-11-13
GPQA Diamond85%mediumEpoch AI2025-11-17
GPQA Diamond66.7%noneEpoch AI2026-08-07
Humanity's Last Exam23.7%#16 of 41, top 40%Epoch AI
Humanity's Last Exam6.8%noneEpoch AI
SimpleQA Verified48%#30 of 77, top 39%highEpoch AI2026-08-27
MMLU-Pro57.9%#46 of 58, top 80%HELM Capabilities
Vectara Hallucination Rate (lower is better)12.1%Vectara Hallucination Leaderboard
Vectara Hallucination Rate (lower is better)10.9%#66 of 96, top 69%Vectara Hallucination Leaderboard
GPQA (HELM)44.2%#37 of 57, top 65%HELM Capabilities
LMArena Expert1470#49 of 273, top 18%highLMArena2026-10-08

Multimodal

GPT-5.1 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1250#55 of 122, top 46%highLMArena2026-10-09
VPCT58.7%#5 of 24, top 21%highEpoch AI
VPCT53.3%mediumEpoch AI
LMArena Document1403#37 of 38, top 98%LMArena2026-09-13

Multilingual

GPT-5.1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1431#55 of 297, top 19%highLMArena2026-10-08
LMArena Chinese1495#49 of 285, top 18%highLMArena2026-10-08
LMArena French1450#69 of 223, top 31%highLMArena2026-10-08
LMArena German1438#60 of 231, top 26%highLMArena2026-10-08
LMArena Japanese1453#23 of 211, top 11%highLMArena2026-10-08
LMArena Korean1401#52 of 213, top 25%highLMArena2026-10-08
LMArena Russian1435#60 of 283, top 22%highLMArena2026-10-08
LMArena Spanish1433#79 of 226, top 35%highLMArena2026-10-08

Instruction Following

GPT-5.1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following93.3%Best of 39highEpoch AI
IFEval93.5%#3 of 57, top 6%HELM Capabilities
LMArena Instruction Following1443#47 of 298, top 16%highLMArena2026-10-08

Long Context

GPT-5.1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench21.1%Epoch AI
CL-bench23.7%#2 of 19, top 11%highEpoch AI
CL-bench Life17.3%#3 of 13, top 24%highEpoch AI
LMArena Longer Query1447#55 of 291, top 19%highLMArena2026-10-08

Writing & Preference

GPT-5.1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1443#56 of 297, top 19%highLMArena2026-10-08
LMArena Creative Writing1427#50 of 295, top 17%highLMArena2026-10-08
WildBench86.3%#2 of 57, top 4%HELM Capabilities
LMArena Multi-Turn1450#54 of 295, top 19%highLMArena2026-10-08
LiveBench Language80.2%Best of 39highEpoch AI

API pricing by provider

GPT-5.1 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.25$10$0.132026-10-10
openai$1.25$10$0.132026-10-10
openrouter$1.25$10$0.132026-10-10

Compare GPT-5.1

Other OpenAI models

Frequently asked questions

How good is GPT-5.1?

GPT-5.1 by OpenAI ranks 53rd of 354 ranked models on the Noometry Index as of October 2026, with a score of 49.0. Its strongest category is instruction following, where it ranks 1st. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

How much does GPT-5.1 cost?

GPT-5.1 costs $1.25 per million input tokens and $10 per million output tokens on OpenAI's own API, with cached input at $0.13.

What is GPT-5.1's context window?

GPT-5.1 accepts up to 400K tokens of input and can write up to 128K tokens in one response.

Is GPT-5.1 open source?

No. GPT-5.1 is proprietary and available only through OpenAI's API and partner platforms.

What are GPT-5.1's strengths and weaknesses?

Relative to other ranked models, GPT-5.1 places best in instruction following, long context, multimodal and lowest in agentic & tool use, knowledge, coding.

What is GPT-5.1 best at?

Its best category is instruction following, where it ranks 1st on Noometry.