Model comparison

GPT-4.1 vs o3

o3 is the stronger model overall, scoring 47.5 to 35.9 on the Noometry Index.

Last verified . 51 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 51 benchmarks with published results for both. GPT-4.1 scores higher in 1 category and o3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where o3 leads 50.2 to 22.3.
  • The biggest single-benchmark swing is ARC-AGI-1: 5.5% for GPT-4.1 and 60.8% for o3.
  • Both cost about the same: $2 input and $8 output per million tokens.
  • GPT-4.1 accepts more context: 1.05M tokens versus 200K.

Side by side

GPT-4.1 and o3 specifications
GPT-4.1o3
ProviderOpenAIOpenAI
Noometry Index35.947.5
Released2025-04-142025-04-16
WeightsProprietaryProprietary
Context window1.05M200K
Max output33K100K
Input $ / M tokens$2$2
Output $ / M tokens$8$8
Results tracked5263

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

GPT-4.1: 34.4 (#238), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGPT-4.1o3
SWE-bench Verified48.5%62.3%
SWE-bench Verified (bash only)39.6%58.4%
Aider Polyglot52.4%81.3%
WeirdML39%52.4%
LMArena Coding13911408
CadEval42%74%
ALE-Bench558.1933.55
GSO—8.8%

Agentic & Tool Use Too close to call

GPT-4.1: 34.7 (#43), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1o3
Berkeley Function Calling Leaderboard54%63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

GPT-4.1: 11.7 (#339), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGPT-4.1o3
ARC-AGI-20.4%6.5%
SimpleBench27%53.1%
Kagi LLM Benchmark52.3%67.6%
ARC-AGI-15.5%60.8%
Chess Puzzles6%38%
EnigmaEval2.2%13.1%
LMArena Hard Prompts13841402
DTBench68.3%84.8%
LMCA25.6%39.7%
Epoch Capabilities Index136.78146.86
ForecastBench61.562.5
CritPt—1.4%
Mystery Game Puzzles—29%

Math o3 leads

GPT-4.1: 22.3 (#280), o3: 50.2 (#58)

Math benchmarks
BenchmarkGPT-4.1o3
FrontierMath (Tiers 1-3)6%33.3%
OTIS Mock AIME 2024-202538.3%84.4%
Omni-MATH47.1%71.4%
LMArena Math13701426
MATH Level 583%97.8%
FrontierMath (Feb 2025 set)5.5%18.7%
FrontierMath Tier 4 (v1)0%2.1%

Knowledge o3 leads

GPT-4.1: 37.1 (#160), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGPT-4.1o3
GPQA Diamond66.9%81.8%
Humanity's Last Exam5.4%20.3%
SimpleQA Verified31.1%49.4%
MMLU-Pro81.1%85.9%
GPQA (HELM)65.9%75.3%
LMArena Expert13641402
Confabulations—14.4%
Vectara Hallucination Rate5.6%—

Multimodal o3 leads

GPT-4.1: 38.2 (#67), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGPT-4.1o3
LMArena Vision12111214
GeoBench72%74%
VPCT—52%

Multilingual o3 leads

GPT-4.1: 49.4 (#133), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGPT-4.1o3
LMArena Non-English13701401
LMArena Chinese13821437
LMArena French13821430
LMArena German13811420
LMArena Japanese13191403
LMArena Korean13391370
LMArena Russian13771406
LMArena Spanish13761395

Instruction Following o3 leads

GPT-4.1: 71.3 (#153), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGPT-4.1o3
IFEval83.8%86.9%
LMArena Instruction Following13671368

Long Context o3 leads

GPT-4.1: 40.0 (#163), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGPT-4.1o3
Fiction.LiveBench63.9%88.9%
LMArena Longer Query13851372
CL-bench—17.8%

Writing & Preference o3 leads

GPT-4.1: 57.6 (#125), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGPT-4.1o3
LMArena Text13831410
LMArena Creative Writing13631359
EQ-Bench Creative Writing14201676
WildBench85.4%86.1%
LMArena Multi-Turn13981405
Short-Story Creative Writing—83.9%

Frequently asked questions

Is GPT-4.1 better than o3?

o3 is the stronger model overall, scoring 47.5 to 35.9 on the Noometry Index.

Which is cheaper, GPT-4.1 or o3?

o3 is cheaper. It lists at $2 per million input tokens and $8 per million output tokens; GPT-4.1 lists at $2 and $8.

Is GPT-4.1 or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 200K.

How many benchmarks do GPT-4.1 and o3 share?

51 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper