OpenAI, proprietary

GPT-4.1

GPT-4.1 by OpenAI ranks 219th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.9. Its strongest category is agentic & tool use, where it ranks 43rd. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 1.05M-token context window.

Last verified

Specifications

Noometry rank
#219 of 354
Index score
35.9
Evidence
Confirmed 52 results
Provider
OpenAI
Released
April 14, 2025
Weights
Proprietary
Reasoning
No
Context window
1.05M
Max output
33K
Input price
$2 / M
Output price
$8 / M
Blended price
$3.50 / M
Output speed
116 tokens/s Kagi
Value
#184 of 219
Knowledge cutoff
April 2024
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

GPT-4.1 category scores
  1. Coding 34.4
  2. Agentic & Tool Use 34.7
  3. Reasoning 11.7
  4. Math 22.3
  5. Knowledge 37.1
  6. Multimodal 38.2
  7. Multilingual 49.4
  8. Instruction Following 71.3
  9. Long Context 40.0
  10. Writing & Preference 57.6
GPT-4.1 category ranks
CategoryScoreRankResults
Coding34.4#2386
Agentic & Tool Use34.7#431
Reasoning11.7#3399
Math22.3#2805
Knowledge37.1#1607
Multimodal38.2#672
Multilingual49.4#1331
Instruction Following71.3#1532
Long Context40.0#1632
Writing & Preference57.6#1255

Strengths and weaknesses

Categories where GPT-4.1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-4.1: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use34.7+4.3#43 of 154, top 28%
Writing & Preference57.6+3.8#125 of 312, top 41%
Multilingual49.4+2.0#133 of 297, top 45%

Weakest categories

GPT-4.1: weakest categories
CategoryScorevs medianRank
Reasoning11.7−11.9#339 of 350, top 97%
Math22.3−14.3#280 of 327, top 86%
Coding34.4−4.4#238 of 340, top 70%

Closest competitors

The models ranked just above and below GPT-4.1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4.1
ModelRankScoreBlended $/MSpeed
Command A#21536.5$4.3828Compare
Grok Build 0.1#21636.4$1.25—Compare
gpt-oss-120b#21736.3$0.070355Compare
Mistral Medium#21836.3$368Compare
Deepseek Coder v2#22035.9——Compare
C4ai Aya Expanse 32b#22135.9——Compare
Llama 3.1 Nemotron 51b Instruct#22235.9——Compare
Nemotron 4 340b Instruct#22335.9——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-4.1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified48.5%#31 of 32, top 97%Epoch AI2026-02-08
SWE-bench Verified (bash only)39.6%#30 of 39, top 77%SWE-bench2025-07-26
Aider Polyglot52.4%#21 of 44, top 48%Epoch AI
WeirdML39%#81 of 119, top 69%Epoch AI
LMArena Coding1391#142 of 294, top 49%LMArena2026-10-08
CadEval42%#8 of 14, top 58%Epoch AI
ALE-Bench558.1#84 of 105, top 80%Epoch AI

Agentic & Tool Use

GPT-4.1 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard54%#15 of 49, top 31%fcBerkeley Function Calling Leaderboard

Reasoning

GPT-4.1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20.4%#73 of 83, top 88%Epoch AI
SimpleBench27%#64 of 77, top 84%Epoch AI
Kagi LLM Benchmark52.3%#59 of 99, top 60%Kagi LLM Benchmark
ARC-AGI-15.5%#76 of 83, top 92%Epoch AI
Chess Puzzles6%#90 of 129, top 70%Epoch AI2026-08-07
EnigmaEval2.2%#30 of 38, top 79%Epoch AI
LMArena Hard Prompts1384#135 of 297, top 46%LMArena2026-10-08
DTBench68.3%#96 of 151, top 64%Epoch AI
LMCA25.6%#85 of 125, top 68%Epoch AI
Epoch Capabilities Index136.78#119 of 213, top 56%Epoch AI2025-04-14
ForecastBench61.5#8 of 72, top 12%Epoch AI

Math

GPT-4.1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)6%#77 of 81, top 96%Epoch AI2026-08-27
OTIS Mock AIME 2024-202538.3%#116 of 173, top 68%Epoch AI2025-04-14
Omni-MATH47.1%#18 of 57, top 32%HELM Capabilities
LMArena Math1370#150 of 285, top 53%LMArena2026-10-08
MATH Level 583%#23 of 79, top 30%Epoch AI2025-04-14
FrontierMath (Feb 2025 set)5.5%#45 of 68, top 67%Epoch AI2025-04-14
FrontierMath Tier 4 (v1)0%#50 of 55, top 91%Epoch AI2025-07-01

Knowledge

GPT-4.1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond66.9%#103 of 186, top 56%Epoch AI2025-04-14
Humanity's Last Exam5.4%#35 of 41, top 86%Epoch AI
SimpleQA Verified31.1%#56 of 77, top 73%Epoch AI2026-08-31
MMLU-Pro81.1%#14 of 58, top 25%HELM Capabilities
Vectara Hallucination Rate (lower is better)5.6%#18 of 96, top 19%Vectara Hallucination Leaderboard
GPQA (HELM)65.9%#16 of 57, top 29%HELM Capabilities
LMArena Expert1364#144 of 273, top 53%LMArena2026-10-08

Multimodal

GPT-4.1 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1211#75 of 122, top 62%LMArena2026-10-09
GeoBench72%#11 of 25, top 44%Epoch AI

Multilingual

GPT-4.1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1370#133 of 297, top 45%LMArena2026-10-08
LMArena Chinese1382#148 of 285, top 52%LMArena2026-10-08
LMArena French1382#130 of 223, top 59%LMArena2026-10-08
LMArena German1381#110 of 231, top 48%LMArena2026-10-08
LMArena Japanese1319#113 of 211, top 54%LMArena2026-10-08
LMArena Korean1339#110 of 213, top 52%LMArena2026-10-08
LMArena Russian1377#130 of 283, top 46%LMArena2026-10-08
LMArena Spanish1376#128 of 226, top 57%LMArena2026-10-08

Instruction Following

GPT-4.1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval83.8%#25 of 57, top 44%HELM Capabilities
LMArena Instruction Following1367#133 of 298, top 45%LMArena2026-10-08

Long Context

GPT-4.1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench63.9%#24 of 47, top 52%Epoch AI
LMArena Longer Query1385#129 of 291, top 45%LMArena2026-10-08

Writing & Preference

GPT-4.1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1383#134 of 297, top 46%LMArena2026-10-08
LMArena Creative Writing1363#117 of 295, top 40%LMArena2026-10-08
EQ-Bench Creative Writing1420#67 of 115, top 59%EQ-Bench
WildBench85.4%#10 of 57, top 18%HELM Capabilities
LMArena Multi-Turn1398#121 of 295, top 42%LMArena2026-10-08

API pricing by provider

GPT-4.1 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$2$8$0.502026-10-10
openai$2$8$0.502026-10-10
openrouter$2$8$0.502026-10-10

Compare GPT-4.1

Other OpenAI models

Frequently asked questions

How good is GPT-4.1?

GPT-4.1 by OpenAI ranks 219th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.9. Its strongest category is agentic & tool use, where it ranks 43rd. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 1.05M-token context window.

How much does GPT-4.1 cost?

GPT-4.1 costs $2 per million input tokens and $8 per million output tokens on OpenAI's own API, with cached input at $0.50.

What is GPT-4.1's context window?

GPT-4.1 accepts up to 1.05M tokens of input and can write up to 33K tokens in one response.

Is GPT-4.1 open source?

No. GPT-4.1 is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-4.1?

GPT-4.1 generated about 116 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-4.1's strengths and weaknesses?

Relative to other ranked models, GPT-4.1 places best in agentic & tool use, writing & preference, multilingual and lowest in reasoning, math, coding.

What is GPT-4.1 best at?

Its best category is agentic & tool use, where it ranks 43rd on Noometry.