OpenAI, proprietary

GPT-5.1-Codex

GPT-5.1-Codex by OpenAI ranks 186th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is agentic & tool use, where it ranks 33rd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

Last verified

Specifications

Noometry rank
#186 of 354
Index score
38.6
Evidence
Reported 6 results
Provider
OpenAI
Released
November 12, 2025
Weights
Proprietary
Reasoning
Yes
Context window
400K
Max output
128K
Input price
$1.25 / M
Output price
$10 / M
Blended price
$3.44 / M
Output speed
Not measured
Value
#178 of 219
Knowledge cutoff
September 2024
Input
text, image, audio

Category scores

Each category score combines every public result we have in that category.

GPT-5.1-Codex category scores
  1. Coding 41.9
  2. Agentic & Tool Use 38.0
  3. Math 30.3
GPT-5.1-Codex category ranks
CategoryScoreRankResults
Coding41.9#1162
Agentic & Tool Use38.0#331
Math30.3#2351

Strengths and weaknesses

Categories where GPT-5.1-Codex places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5.1-Codex: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use38.0+7.6#33 of 154, top 22%

Weakest categories

GPT-5.1-Codex: weakest categories
CategoryScorevs medianRank
Math30.3−6.2#235 of 327, top 72%

Closest competitors

The models ranked just above and below GPT-5.1-Codex. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.1-Codex
ModelRankScoreBlended $/MSpeed
Qwen3.6 Flash#18238.8$0.42—Compare
Olmo 3 32b Think#18338.7——Compare
Hunyuan Large 2025 02 10#18438.6——Compare
Trinity Large Thinking#18538.6$0.39—Compare
Sonar#18738.5$1—Compare
MiniMax-M2.5#18838.3$0.5249Compare
Nova Premier 1.0#18938.3$59Compare
Qwen3-Coder 480B-A35B Instruct#19038.1$367Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5.1-Codex Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)66%#15 of 39, top 39%mediumSWE-bench2025-11-24
LMArena WebDev1337#92 of 113, top 82%LMArena2026-10-08
ALE-Bench1,209Epoch AI
ALE-Bench1,245#29 of 105, top 28%Epoch AI

Agentic & Tool Use

GPT-5.1-Codex Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench57.8%Epoch AI
Terminal-Bench60.4%#13 of 41, top 32%Epoch AI
METR Time Horizons70.8%#9 of 32, top 29%Epoch AI

Math

GPT-5.1-Codex Math benchmark results
BenchmarkScorePositionSettingSourceDate
ProofBench9%#63 of 77, top 82%Epoch AI

API pricing by provider

GPT-5.1-Codex API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.25$10$0.132026-10-10
openrouter$1.25$10$0.132026-10-10

Compare GPT-5.1-Codex

Other OpenAI models

Frequently asked questions

How good is GPT-5.1-Codex?

GPT-5.1-Codex by OpenAI ranks 186th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is agentic & tool use, where it ranks 33rd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

How much does GPT-5.1-Codex cost?

GPT-5.1-Codex costs $1.25 per million input tokens and $10 per million output tokens on azure, with cached input at $0.13.

What is GPT-5.1-Codex's context window?

GPT-5.1-Codex accepts up to 400K tokens of input and can write up to 128K tokens in one response.

Is GPT-5.1-Codex open source?

No. GPT-5.1-Codex is proprietary and available only through OpenAI's API and partner platforms.

What are GPT-5.1-Codex's strengths and weaknesses?

Relative to other ranked models, GPT-5.1-Codex places best in agentic & tool use and lowest in math.

What is GPT-5.1-Codex best at?

Its best category is agentic & tool use, where it ranks 33rd on Noometry.