OpenAI, proprietary

GPT-5.2

GPT-5.2 by OpenAI ranks 34th of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.1. Its strongest category is multimodal, where it ranks 7th. API pricing starts at $1.75 per million input tokens and $14 per million output tokens, with a 400K-token context window.

Last verified

Specifications

Noometry rank
#34 of 354
Index score
54.1
Evidence
Confirmed 67 results
Provider
OpenAI
Released
December 11, 2025
Weights
Proprietary
Reasoning
Yes
Context window
400K
Max output
128K
Input price
$1.75 / M
Output price
$14 / M
Blended price
$4.81 / M
Output speed
15 tokens/s Kagi
Value
#179 of 219
Knowledge cutoff
August 2025
Input
text, image

Category scores

Each category score combines every public result we have in that category.

GPT-5.2 category scores
  1. Coding 51.6
  2. Agentic & Tool Use 40.2
  3. Reasoning 50.2
  4. Math 60.0
  5. Knowledge 59.3
  6. Multimodal 51.3
  7. Multilingual 53.4
  8. Instruction Following 74.7
  9. Long Context 44.0
  10. Writing & Preference 66.8
GPT-5.2 category ranks
CategoryScoreRankResults
Coding51.6#377
Agentic & Tool Use40.2#249
Reasoning50.2#3512
Math60.0#386
Knowledge59.3#325
Multimodal51.3#73
Multilingual53.4#671
Instruction Following74.7#891
Long Context44.0#782
Writing & Preference66.8#324

Strengths and weaknesses

Categories where GPT-5.2 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5.2: strongest categories
CategoryScorevs medianRank
Multimodal51.3+12.7#7 of 128, top 6%
Reasoning50.2+26.6#35 of 350, top 10%
Knowledge59.3+21.9#32 of 314, top 11%

Weakest categories

GPT-5.2: weakest categories
CategoryScorevs medianRank
Instruction Following74.7+3.4#89 of 305, top 30%
Long Context44.0+3.1#78 of 296, top 27%
Multilingual53.4+6.0#67 of 297, top 23%

Closest competitors

The models ranked just above and below GPT-5.2. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.2
ModelRankScoreBlended $/MSpeed
GPT-5.6 Luna#3054.6$0.4512Compare
DeepSeek V4 Pro#3154.3$0.9916Compare
Gemini 3.5 Flash#3254.2$3.38—Compare
Gemini 3.6 Flash#3354.1$1.50—Compare
DeepSeek V4 Flash#3553.6$0.266Compare
GPT-6 Luna#3653.3$0.20—Compare
Grok 4.7#3753.1$3—Compare
DeepSeek V4.1 Flash#3852.8$0.26—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5.2 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified73.8%#17 of 32, top 54%highEpoch AI2026-02-12
SWE-bench Verified (bash only)72.8%#7 of 39, top 18%highSWE-bench2026-02-17
SWE-bench Verified (bash only)71.8%highSWE-bench2025-12-11
LMArena WebDev1416#66 of 113, top 59%LMArena2026-10-08
SWE-bench Multilingual66.7%#9 of 13, top 70%highSWE-bench2026-02-13
GSO27.4%#11 of 31, top 36%highEpoch AI
WeirdML49.6%lowEpoch AI
WeirdML63.4%mediumEpoch AI
WeirdML49.6%noneEpoch AI
WeirdML72.2%#16 of 119, top 14%xhighEpoch AI
LMArena Coding1447#90 of 294, top 31%LMArena2026-10-08
LMArena Coding1442highLMArena2026-10-08
ALE-Bench1,294#27 of 105, top 26%highEpoch AI
ALE-Bench1,250mediumEpoch AI
AlgoTune2.05Best of 18mediumEpoch AI

Agentic & Tool Use

GPT-5.2 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench64.9%#9 of 41, top 22%Epoch AI
Terminal-Bench64.9%mediumEpoch AI
Berkeley Function Calling Leaderboard55.9%#13 of 49, top 27%fcBerkeley Function Calling Leaderboard
GDPval49.7%Best of 11noneEpoch AI
Remote Labor Index2.1%Epoch AI
Remote Labor Index2.5%#10 of 14, top 72%mediumEpoch AI
τ²-bench Airline83%#2 of 7, top 29%highτ²-bench2026-02-26
τ²-bench Banking32.2%#13 of 26, top 50%highτ²-bench2026-02-26
τ²-bench Retail81.6%#2 of 7, top 29%highτ²-bench2026-02-26
τ²-bench Telecom89.7%#5 of 7, top 72%highτ²-bench2026-02-26
DeepResearch Bench41.1%#20 of 24, top 84%lowEpoch AI
LMArena Search1172LMArena2026-08-24
LMArena Search1207#11 of 32, top 35%LMArena2026-08-24
METR Time Horizons75.3%#4 of 32, top 13%highEpoch AI
Vending-Bench 23,591#40 of 60, top 67%Epoch AI

Reasoning

GPT-5.2 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20.8%Epoch AI
ARC-AGI-243.3%highEpoch AI
ARC-AGI-29.7%lowEpoch AI
ARC-AGI-226.7%mediumEpoch AI
ARC-AGI-252.9%#33 of 83, top 40%xhighEpoch AI
SimpleBench45.8%#47 of 77, top 62%Epoch AI
SimpleBench45.8%highEpoch AI
Kagi LLM Benchmark73.3%#17 of 99, top 18%Kagi LLM Benchmark
NYT Connections (extended)83.6%#31 of 91, top 35%xhigh reasoningLech Mazur benchmarks
ARC-AGI-112.3%Epoch AI
ARC-AGI-178.7%highEpoch AI
ARC-AGI-155.7%lowEpoch AI
ARC-AGI-172.7%mediumEpoch AI
ARC-AGI-112.3%noneEpoch AI
ARC-AGI-186.2%#35 of 83, top 43%xhighEpoch AI
Chess Puzzles40%highEpoch AI2025-12-11
Chess Puzzles23%lowEpoch AI2025-12-11
Chess Puzzles40%mediumEpoch AI2025-12-11
Chess Puzzles4%noneEpoch AI2026-07-13
Chess Puzzles49%#11 of 129, top 9%xhighEpoch AI2025-12-15
EnigmaEval10.4%#14 of 38, top 37%Epoch AI
EBR-Bench23%#13 of 24, top 55%xhighEpoch AI2026-06-26
LMArena Hard Prompts1445#72 of 297, top 25%LMArena2026-10-08
LMArena Hard Prompts1428highLMArena2026-10-08
Mystery Game Puzzles23%#36 of 74, top 49%highEpoch AI2026-08-06
Mystery Game Puzzles10%lowEpoch AI2026-08-27
Mystery Game Puzzles22%mediumEpoch AI2026-08-28
Mystery Game Puzzles14%noneEpoch AI2026-08-27
DTBench90.9%#33 of 151, top 22%xhighEpoch AI
LMCA43.9%#38 of 125, top 31%xhighEpoch AI
Epoch Capabilities Index153.45#41 of 213, top 20%Epoch AI2025-12-11
ForecastBench60.1#31 of 72, top 44%Epoch AI

Math

GPT-5.2 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)67.4%#27 of 81, top 34%xhighEpoch AI2026-06-11
FrontierMath Tier 431.7%#28 of 63, top 45%xhighEpoch AI2026-06-11
MathArena Final-Answer Competitions72%#12 of 29, top 42%highMathArena
OTIS Mock AIME 2024-202596.1%#29 of 173, top 17%highEpoch AI2025-12-11
OTIS Mock AIME 2024-202578.9%lowEpoch AI2025-12-11
OTIS Mock AIME 2024-202593.9%mediumEpoch AI2025-12-11
OTIS Mock AIME 2024-202562.2%noneEpoch AI2026-07-13
OTIS Mock AIME 2024-202596.1%xhighEpoch AI2025-12-13
ProofBench15%#58 of 77, top 76%xhighEpoch AI
LMArena Math1435LMArena2026-10-08
LMArena Math1440#72 of 285, top 26%highLMArena2026-10-08
FrontierMath (Feb 2025 set)40.3%highEpoch AI2025-12-11
FrontierMath (Feb 2025 set)26.6%lowEpoch AI2025-12-11
FrontierMath (Feb 2025 set)36.9%mediumEpoch AI2025-12-11
FrontierMath (Feb 2025 set)40.7%#8 of 68, top 12%xhighEpoch AI2025-12-13
FrontierMath Tier 4 (v1)18.8%#10 of 55, top 19%highEpoch AI2025-12-11
FrontierMath Tier 4 (v1)6.3%lowEpoch AI2025-12-11
FrontierMath Tier 4 (v1)16.7%mediumEpoch AI2025-12-11
FrontierMath Tier 4 (v1)18.8%xhighEpoch AI2025-12-14

Knowledge

GPT-5.2 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond88.2%highEpoch AI2025-12-11
GPQA Diamond82.7%lowEpoch AI2025-12-11
GPQA Diamond87.9%mediumEpoch AI2025-12-11
GPQA Diamond73.2%noneEpoch AI2026-07-13
GPQA Diamond91.4%#26 of 186, top 14%xhighEpoch AI2025-12-13
Humanity's Last Exam27.8%#12 of 41, top 30%Epoch AI
SimpleQA Verified34.3%highEpoch AI2026-08-27
SimpleQA Verified32.8%lowEpoch AI2026-08-27
SimpleQA Verified32.7%mediumEpoch AI2026-08-27
SimpleQA Verified37.1%#45 of 77, top 59%xhighEpoch AI2026-08-27
Vectara Hallucination Rate (lower is better)10.8%Vectara Hallucination Leaderboard
Vectara Hallucination Rate (lower is better)8.4%#39 of 96, top 41%Vectara Hallucination Leaderboard
LMArena Expert1438LMArena2026-10-08
LMArena Expert1445#74 of 273, top 28%highLMArena2026-10-08

Multimodal

GPT-5.2 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1268#39 of 122, top 32%LMArena2026-10-09
LMArena Vision1246highLMArena2026-10-09
VPCT67%highEpoch AI
VPCT84%#2 of 24, top 9%xhighEpoch AI
Furniture Assembly38.3%#16 of 31, top 52%xhighEpoch AI2026-09-10
LMArena Document1405#36 of 38, top 95%highLMArena2026-09-13

Multilingual

GPT-5.2 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1425#67 of 297, top 23%LMArena2026-10-08
LMArena Non-English1405LMArena2026-10-08
LMArena Chinese1460#91 of 285, top 32%LMArena2026-10-08
LMArena Chinese1447highLMArena2026-10-08
LMArena French1438LMArena2026-10-08
LMArena French1455#61 of 223, top 28%highLMArena2026-10-08
LMArena German1448#46 of 231, top 20%LMArena2026-10-08
LMArena German1444highLMArena2026-10-08
LMArena Japanese1418LMArena2026-10-08
LMArena Japanese1420#41 of 211, top 20%highLMArena2026-10-08
LMArena Korean1392#66 of 213, top 31%LMArena2026-10-08
LMArena Korean1374LMArena2026-10-08
LMArena Russian1415LMArena2026-10-08
LMArena Russian1440#52 of 283, top 19%LMArena2026-10-08
LMArena Spanish1433#78 of 226, top 35%LMArena2026-10-08
LMArena Spanish1413highLMArena2026-10-08

Instruction Following

GPT-5.2 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1417#78 of 298, top 27%LMArena2026-10-08
LMArena Instruction Following1409highLMArena2026-10-08

Long Context

GPT-5.2 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench18.2%#11 of 19, top 58%Epoch AI
CL-bench18.1%highEpoch AI
LMArena Longer Query1428#80 of 291, top 28%LMArena2026-10-08
LMArena Longer Query1413LMArena2026-10-08

Writing & Preference

GPT-5.2 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1439#65 of 297, top 22%LMArena2026-10-08
LMArena Text1416highLMArena2026-10-08
LMArena Creative Writing1401#77 of 295, top 27%LMArena2026-10-08
LMArena Creative Writing1376highLMArena2026-10-08
EQ-Bench Creative Writing1703#30 of 115, top 27%EQ-Bench
LMArena Multi-Turn1458#40 of 295, top 14%LMArena2026-10-08
LMArena Multi-Turn1419highLMArena2026-10-08

API pricing by provider

GPT-5.2 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.75$14$0.132026-10-10
openai$1.75$14$0.172026-10-10
openrouter$1.75$14$0.172026-10-10

Compare GPT-5.2

Other OpenAI models

Frequently asked questions

How good is GPT-5.2?

GPT-5.2 by OpenAI ranks 34th of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.1. Its strongest category is multimodal, where it ranks 7th. API pricing starts at $1.75 per million input tokens and $14 per million output tokens, with a 400K-token context window.

How much does GPT-5.2 cost?

GPT-5.2 costs $1.75 per million input tokens and $14 per million output tokens on OpenAI's own API, with cached input at $0.17.

What is GPT-5.2's context window?

GPT-5.2 accepts up to 400K tokens of input and can write up to 128K tokens in one response.

Is GPT-5.2 open source?

No. GPT-5.2 is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-5.2?

GPT-5.2 generated about 15 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-5.2's strengths and weaknesses?

Relative to other ranked models, GPT-5.2 places best in multimodal, reasoning, knowledge and lowest in instruction following, long context, multilingual.

What is GPT-5.2 best at?

Its best category is multimodal, where it ranks 7th on Noometry.