OpenAI, proprietary
GPT-4.1 nano
GPT-4.1 nano by OpenAI ranks 327th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.9. Its strongest category is agentic & tool use, where it ranks 104th. API pricing starts at $0.10 per million input tokens and $0.40 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #327 of 354
- Index score
- 27.9
- Evidence
- Confirmed 38 results
- Provider
- OpenAI
- Released
- April 14, 2025
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 1.05M
- Max output
- 33K
- Input price
- $0.10 / M
- Output price
- $0.40 / M
- Blended price
- $0.18 / M
- Output speed
- 135 tokens/s Kagi
- Value
- #48 of 219
- Knowledge cutoff
- April 2024
- Input
- text, image
Category scores
Each category score combines every public result we have in that category.
- Coding 24.1
- Agentic & Tool Use 26.5
- Reasoning 8.5
- Math 26.9
- Knowledge 21.8
- Multimodal 29.2
- Multilingual 41.6
- Instruction Following 67.8
- Long Context 23.7
- Writing & Preference 40.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 24.1 | #330 | 4 |
| Agentic & Tool Use | 26.5 | #104 | 1 |
| Reasoning | 8.5 | #349 | 7 |
| Math | 26.9 | #252 | 4 |
| Knowledge | 21.8 | #273 | 5 |
| Multimodal | 29.2 | #113 | 1 |
| Multilingual | 41.6 | #205 | 1 |
| Instruction Following | 67.8 | #193 | 2 |
| Long Context | 23.7 | #296 | 2 |
| Writing & Preference | 40.5 | #243 | 5 |
Strengths and weaknesses
Categories where GPT-4.1 nano places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 67.8 | −3.5 | #193 of 305, top 64% |
| Agentic & Tool Use | 26.5 | −3.8 | #104 of 154, top 68% |
| Multilingual | 41.6 | −5.8 | #205 of 297, top 70% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 23.7 | −17.2 | #296 of 296, top 100% |
| Reasoning | 8.5 | −15.1 | #349 of 350, top 100% |
| Coding | 24.1 | −14.6 | #330 of 340, top 98% |
Closest competitors
The models ranked just above and below GPT-4.1 nano. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Llama 3-70B | #323 | 28.8 | — | 104 | Compare |
| GPT-4o | #324 | 28.6 | $4.38 | — | Compare |
| Ministral 8B | #325 | 28.2 | $0.15 | — | Compare |
| Gemma 3 4B | #326 | 28.1 | $0.05 | 72 | Compare |
| Phi 3 Mini 4k Instruct | #328 | 27.9 | — | — | Compare |
| Yi-34B | #329 | 27.8 | — | — | Compare |
| Llama 4 Scout | #330 | 27.7 | $0.15 | 272 | Compare |
| Llama 3.2 90B | #331 | 27.5 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 8.9% | #42 of 44, top 96% | Epoch AI | ||
| SciCode | 25.9% | #112 of 121, top 93% | Epoch AI | ||
| WeirdML | 19% | #106 of 119, top 90% | Epoch AI | ||
| LMArena Coding | 1306 | #197 of 294, top 68% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 33% | #30 of 49, top 62% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0% | #76 of 83, top 92% | Epoch AI | ||
| Kagi LLM Benchmark | 33.3% | #90 of 99, top 91% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 0% | #83 of 83, top 100% | Epoch AI | ||
| CritPt | 0% | #111 of 134, top 83% | Epoch AI | ||
| LMArena Hard Prompts | 1286 | #196 of 297, top 66% | LMArena | 2026-10-08 | |
| DTBench | 52.5% | #130 of 151, top 87% | Epoch AI | ||
| LMCA | 5.5% | #121 of 125, top 97% | Epoch AI | ||
| Epoch Capabilities Index | 129.62 | #140 of 213, top 66% | Epoch AI | 2025-04-14 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 28.9% | #122 of 173, top 71% | Epoch AI | 2025-04-14 | |
| Omni-MATH | 36.7% | #31 of 57, top 55% | HELM Capabilities | ||
| LMArena Math | 1274 | #195 of 285, top 69% | LMArena | 2026-10-08 | |
| MATH Level 5 | 70% | #31 of 79, top 40% | Epoch AI | 2025-04-14 | |
| FrontierMath (Feb 2025 set) | 1% | #59 of 68, top 87% | Epoch AI | 2025-04-14 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 48.9% | #131 of 186, top 71% | Epoch AI | 2025-04-14 | |
| SimpleQA Verified | 6% | #77 of 77, top 100% | Epoch AI | 2026-08-31 | |
| MMLU-Pro | 55% | #48 of 58, top 83% | HELM Capabilities | ||
| GPQA (HELM) | 50.7% | #33 of 57, top 58% | HELM Capabilities | ||
| LMArena Expert | 1272 | #187 of 273, top 69% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1063 | #113 of 122, top 93% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1260 | #205 of 297, top 70% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1270 | #199 of 285, top 70% | LMArena | 2026-10-08 | |
| LMArena German | 1288 | #154 of 231, top 67% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1198 | #162 of 211, top 77% | LMArena | 2026-10-08 | |
| LMArena Russian | 1261 | #204 of 283, top 73% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 84.3% | #21 of 57, top 37% | HELM Capabilities | ||
| LMArena Instruction Following | 1267 | #199 of 298, top 67% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 25% | #47 of 47, top 100% | Epoch AI | ||
| LMArena Longer Query | 1283 | #201 of 291, top 70% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1285 | #204 of 297, top 69% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1260 | #199 of 295, top 68% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 946 | #98 of 115, top 86% | EQ-Bench | ||
| WildBench | 81.2% | #25 of 57, top 44% | HELM Capabilities | ||
| LMArena Multi-Turn | 1277 | #202 of 295, top 69% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.10 | $0.40 | $0.025 | 2026-10-10 |
| openai | $0.10 | $0.40 | $0.025 | 2026-10-10 |
| openrouter | $0.10 | $0.40 | $0.025 | 2026-10-10 |
Compare GPT-4.1 nano
- GPT-4.1 nano vs Gemma 3 4B
- GPT-4.1 nano vs Phi 3 Mini 4k Instruct
- GPT-4.1 nano vs Ministral 8B
- GPT-4.1 nano vs Yi-34B
- GPT-4.1 nano vs GPT-4o
- GPT-4.1 nano vs Llama 4 Scout
- GPT-4.1 nano vs Claude Fable 5.1
- GPT-4.1 nano vs Gemini 3.8 Flash
- GPT-4.1 nano vs Kimi K3
- GPT-4.1 nano vs Grok 4.6
- GPT-4.1 nano vs Qwen3.8 Max
- GPT-4.1 nano vs GLM-5.3
- GPT-4.1 nano vs Muse Spark 1.3
- GPT-4.1 nano vs DeepSeek V4 Pro
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-4.1 nano?
GPT-4.1 nano by OpenAI ranks 327th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.9. Its strongest category is agentic & tool use, where it ranks 104th. API pricing starts at $0.10 per million input tokens and $0.40 per million output tokens, with a 1.05M-token context window.
How much does GPT-4.1 nano cost?
GPT-4.1 nano costs $0.10 per million input tokens and $0.40 per million output tokens on OpenAI's own API, with cached input at $0.025.
What is GPT-4.1 nano's context window?
GPT-4.1 nano accepts up to 1.05M tokens of input and can write up to 33K tokens in one response.
Is GPT-4.1 nano open source?
No. GPT-4.1 nano is proprietary and available only through OpenAI's API and partner platforms.
How fast is GPT-4.1 nano?
GPT-4.1 nano generated about 135 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are GPT-4.1 nano's strengths and weaknesses?
Relative to other ranked models, GPT-4.1 nano places best in instruction following, agentic & tool use, multilingual and lowest in long context, reasoning, coding.
What is GPT-4.1 nano best at?
Its best category is agentic & tool use, where it ranks 104th on Noometry.