Model comparison

GPT-4.1 vs GPT-4.5

GPT-4.5 is the stronger model overall, scoring 37.2 to 35.9 on the Noometry Index.

Last verified . 31 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GPT-4.1 scores higher in 4 categories and GPT-4.5 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-4.5 leads 32.6 to 22.3.
  • The biggest single-benchmark swing is Aider Polyglot: 52.4% for GPT-4.1 and 44.9% for GPT-4.5.

Side by side

GPT-4.1 and GPT-4.5 specifications
GPT-4.1GPT-4.5
ProviderOpenAIOpenAI
Noometry Index35.937.2
Released2025-04-142025-02-27
WeightsProprietaryProprietary
Context window1.05M—
Max output33K—
Input $ / M tokens$2—
Output $ / M tokens$8—
Results tracked5242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

GPT-4.1: 34.4 (#238), GPT-4.5: 42.2 (#109)

Coding benchmarks
BenchmarkGPT-4.1GPT-4.5
Aider Polyglot52.4%44.9%
WeirdML39%39.4%
LMArena Coding13911396
SWE-bench Verified48.5%—
SWE-bench Verified (bash only)39.6%—
LiveBench Coding—75.2%
CadEval42%—
ALE-Bench558.1—

Agentic & Tool Use GPT-4.1 leads

GPT-4.1: 34.7 (#43), GPT-4.5: 27.9 (#97)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1GPT-4.5
Berkeley Function Calling Leaderboard54%—
Cybench—17.5%

Reasoning GPT-4.5 leads

GPT-4.1: 11.7 (#339), GPT-4.5: 13.9 (#330)

Reasoning benchmarks
BenchmarkGPT-4.1GPT-4.5
ARC-AGI-20.4%0.8%
SimpleBench27%34.5%
ARC-AGI-15.5%10.3%
EnigmaEval2.2%3.2%
LMArena Hard Prompts13841403
Epoch Capabilities Index136.78136.74
ForecastBench61.561.7
Kagi LLM Benchmark52.3%—
Chess Puzzles6%—
LiveBench Reasoning—71.1%
DTBench68.3%—
LiveBench Data Analysis—64.3%
LMCA25.6%—
LiveBench—69%

Math GPT-4.5 leads

GPT-4.1: 22.3 (#280), GPT-4.5: 32.6 (#211)

Knowledge GPT-4.1 leads

GPT-4.1: 37.1 (#160), GPT-4.5: 32.5 (#211)

Knowledge benchmarks
BenchmarkGPT-4.1GPT-4.5
GPQA Diamond66.9%68.7%
Humanity's Last Exam5.4%5.4%
LMArena Expert13641394
SimpleQA Verified31.1%—
MMLU-Pro81.1%—
Confabulations—13.6%
Vectara Hallucination Rate5.6%—
GPQA (HELM)65.9%—

Multimodal Too close to call

GPT-4.1: 38.2 (#67), GPT-4.5: 37.6 (#71)

Multimodal benchmarks
BenchmarkGPT-4.1GPT-4.5
LMArena Vision12111195
GeoBench72%—
VPCT—45%

Multilingual GPT-4.5 leads

GPT-4.1: 49.4 (#133), GPT-4.5: 52.5 (#83)

Multilingual benchmarks
BenchmarkGPT-4.1GPT-4.5
LMArena Non-English13701413
LMArena Chinese13821421
LMArena French13821418
LMArena German13811457
LMArena Japanese13191416
LMArena Korean13391392
LMArena Russian13771419
LMArena Spanish1376—

Instruction Following GPT-4.5 leads

GPT-4.1: 71.3 (#153), GPT-4.5: 72.6 (#134)

Instruction Following benchmarks
BenchmarkGPT-4.1GPT-4.5
LMArena Instruction Following13671404
LiveBench Instruction Following—72.3%
IFEval83.8%—

Long Context Too close to call

GPT-4.1: 40.0 (#163), GPT-4.5: 40.4 (#155)

Long Context benchmarks
BenchmarkGPT-4.1GPT-4.5
Fiction.LiveBench63.9%63.9%
LMArena Longer Query13851406

Writing & Preference Too close to call

GPT-4.1: 57.6 (#125), GPT-4.5: 56.9 (#134)

Writing & Preference benchmarks
BenchmarkGPT-4.1GPT-4.5
LMArena Text13831417
LMArena Creative Writing13631394
EQ-Bench Creative Writing14201258
LMArena Multi-Turn13981444
Short-Story Creative Writing—75.6%
WildBench85.4%—
LiveBench Language—61.5%

Frequently asked questions

Is GPT-4.1 better than GPT-4.5?

GPT-4.5 is the stronger model overall, scoring 37.2 to 35.9 on the Noometry Index.

Is GPT-4.1 or GPT-4.5 better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 34.4 in the Noometry coding category.

How many benchmarks do GPT-4.1 and GPT-4.5 share?

31 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and GPT-4.5 has 42.

Related comparisons

Go deeper