OpenAI, proprietary
GPT-5.1-Codex
GPT-5.1-Codex by OpenAI ranks 186th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is agentic & tool use, where it ranks 33rd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.
Last verified
Specifications
- Noometry rank
- #186 of 354
- Index score
- 38.6
- Evidence
- Reported 6 results
- Provider
- OpenAI
- Released
- November 12, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 400K
- Max output
- 128K
- Input price
- $1.25 / M
- Output price
- $10 / M
- Blended price
- $3.44 / M
- Output speed
- Not measured
- Value
- #178 of 219
- Knowledge cutoff
- September 2024
- Input
- text, image, audio
Category scores
Each category score combines every public result we have in that category.
- Coding 41.9
- Agentic & Tool Use 38.0
- Math 30.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 41.9 | #116 | 2 |
| Agentic & Tool Use | 38.0 | #33 | 1 |
| Math | 30.3 | #235 | 1 |
Strengths and weaknesses
Categories where GPT-5.1-Codex places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 38.0 | +7.6 | #33 of 154, top 22% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 30.3 | −6.2 | #235 of 327, top 72% |
Closest competitors
The models ranked just above and below GPT-5.1-Codex. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen3.6 Flash | #182 | 38.8 | $0.42 | — | Compare |
| Olmo 3 32b Think | #183 | 38.7 | — | — | Compare |
| Hunyuan Large 2025 02 10 | #184 | 38.6 | — | — | Compare |
| Trinity Large Thinking | #185 | 38.6 | $0.39 | — | Compare |
| Sonar | #187 | 38.5 | $1 | — | Compare |
| MiniMax-M2.5 | #188 | 38.3 | $0.52 | 49 | Compare |
| Nova Premier 1.0 | #189 | 38.3 | $5 | 9 | Compare |
| Qwen3-Coder 480B-A35B Instruct | #190 | 38.1 | $3 | 67 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 66% | #15 of 39, top 39% | medium | SWE-bench | 2025-11-24 |
| LMArena WebDev | 1337 | #92 of 113, top 82% | LMArena | 2026-10-08 | |
| ALE-Bench | 1,209 | Epoch AI | |||
| ALE-Bench | 1,245 | #29 of 105, top 28% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 57.8% | Epoch AI | |||
| Terminal-Bench | 60.4% | #13 of 41, top 32% | Epoch AI | ||
| METR Time Horizons | 70.8% | #9 of 32, top 29% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ProofBench | 9% | #63 of 77, top 82% | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $1.25 | $10 | $0.13 | 2026-10-10 |
| openrouter | $1.25 | $10 | $0.13 | 2026-10-10 |
Compare GPT-5.1-Codex
- GPT-5.1-Codex vs GPT-5-Codex
- GPT-5.1-Codex vs Trinity Large Thinking
- GPT-5.1-Codex vs Sonar
- GPT-5.1-Codex vs Hunyuan Large 2025 02 10
- GPT-5.1-Codex vs MiniMax-M2.5
- GPT-5.1-Codex vs Olmo 3 32b Think
- GPT-5.1-Codex vs Nova Premier 1.0
- GPT-5.1-Codex vs Claude Fable 5.1
- GPT-5.1-Codex vs Gemini 3.8 Flash
- GPT-5.1-Codex vs Kimi K3
- GPT-5.1-Codex vs Grok 4.6
- GPT-5.1-Codex vs Qwen3.8 Max
- GPT-5.1-Codex vs GLM-5.3
- GPT-5.1-Codex vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-5.1-Codex?
GPT-5.1-Codex by OpenAI ranks 186th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is agentic & tool use, where it ranks 33rd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.
How much does GPT-5.1-Codex cost?
GPT-5.1-Codex costs $1.25 per million input tokens and $10 per million output tokens on azure, with cached input at $0.13.
What is GPT-5.1-Codex's context window?
GPT-5.1-Codex accepts up to 400K tokens of input and can write up to 128K tokens in one response.
Is GPT-5.1-Codex open source?
No. GPT-5.1-Codex is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-5.1-Codex's strengths and weaknesses?
Relative to other ranked models, GPT-5.1-Codex places best in agentic & tool use and lowest in math.
What is GPT-5.1-Codex best at?
Its best category is agentic & tool use, where it ranks 33rd on Noometry.