Model comparison

Amazon Nova Pro vs GPT-4.1

GPT-4.1 is the stronger model overall, scoring 35.9 to 31.0 on the Noometry Index. Amazon Nova Pro costs 2.5× less per token, which makes it the better buy when GPT-4.1's lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

Amazon Nova Pro Amazon

31.0

Rank #281 Confirmed

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Amazon Nova Pro scores higher in 3 categories and GPT-4.1 in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GPT-4.1 leads 34.7 to 16.7.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 25% for Amazon Nova Pro and 54% for GPT-4.1.
  • Amazon Nova Pro is cheaper at $0.80 / $3.20 per million input/output tokens, against $2 / $8 for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 300K.

Side by side

Amazon Nova Pro and GPT-4.1 specifications
Amazon Nova ProGPT-4.1
ProviderAmazonOpenAI
Noometry Index31.035.9
Released2024-12-032025-04-14
WeightsProprietaryProprietary
Context window300K1.05M
Max output10K33K
Input $ / M tokens$0.80$2
Output $ / M tokens$3.20$8
Results tracked3852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Amazon Nova Pro: 35.1 (#229), GPT-4.1: 34.4 (#238)

Coding benchmarks
BenchmarkAmazon Nova ProGPT-4.1
LMArena Coding12701391
SWE-bench Verified—48.5%
SWE-bench Verified (bash only)—39.6%
Aider Polyglot—52.4%
WeirdML—39%
LiveBench Coding38.1%—
CadEval—42%
ALE-Bench—558.1

Agentic & Tool Use GPT-4.1 leads

Amazon Nova Pro: 16.7 (#147), GPT-4.1: 34.7 (#43)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova ProGPT-4.1
Berkeley Function Calling Leaderboard25%54%
TheAgentCompany1.7%—

Reasoning Amazon Nova Pro leads

Amazon Nova Pro: 20.0 (#243), GPT-4.1: 11.7 (#339)

Reasoning benchmarks
BenchmarkAmazon Nova ProGPT-4.1
LMArena Hard Prompts12461384
Epoch Capabilities Index123.8136.78
ARC-AGI-2—0.4%
SimpleBench—27%
Kagi LLM Benchmark—52.3%
ARC-AGI-1—5.5%
Chess Puzzles—6%
EnigmaEval—2.2%
LiveBench Reasoning32.6%—
DTBench—68.3%
LiveBench Data Analysis48.3%—
LMCA—25.6%
ForecastBench—61.5
LiveBench43.5%—

Math Amazon Nova Pro leads

Amazon Nova Pro: 28.5 (#243), GPT-4.1: 22.3 (#280)

Math benchmarks
BenchmarkAmazon Nova ProGPT-4.1
Omni-MATH24.2%47.1%
LMArena Math12521370
FrontierMath (Tiers 1-3)—6%
OTIS Mock AIME 2024-2025—38.3%
LiveBench Math38%—
MATH Level 5—83%
FrontierMath (Feb 2025 set)—5.5%
FrontierMath Tier 4 (v1)—0%

Knowledge GPT-4.1 leads

Amazon Nova Pro: 27.4 (#250), GPT-4.1: 37.1 (#160)

Knowledge benchmarks
BenchmarkAmazon Nova ProGPT-4.1
Humanity's Last Exam4.4%5.4%
MMLU-Pro67.3%81.1%
Vectara Hallucination Rate5.1%5.6%
GPQA (HELM)44.6%65.9%
LMArena Expert12111364
GPQA Diamond—66.9%
SimpleQA Verified—31.1%
Confabulations30.1%—
MMLU82%—

Multimodal GPT-4.1 leads

Amazon Nova Pro: 25.0 (#126), GPT-4.1: 38.2 (#67)

Multimodal benchmarks
BenchmarkAmazon Nova ProGPT-4.1
LMArena Vision9801211
GeoBench—72%

Multilingual GPT-4.1 leads

Amazon Nova Pro: 39.7 (#223), GPT-4.1: 49.4 (#133)

Multilingual benchmarks
BenchmarkAmazon Nova ProGPT-4.1
LMArena Non-English12341370
LMArena Chinese12441382
LMArena French12711382
LMArena German12431381
LMArena Japanese12001319
LMArena Korean12031339
LMArena Russian12401377
LMArena Spanish11821376

Instruction Following GPT-4.1 leads

Amazon Nova Pro: 64.9 (#226), GPT-4.1: 71.3 (#153)

Instruction Following benchmarks
BenchmarkAmazon Nova ProGPT-4.1
IFEval81.5%83.8%
LMArena Instruction Following12351367
LiveBench Instruction Following67.1%—

Long Context GPT-4.1 leads

Amazon Nova Pro: 38.1 (#205), GPT-4.1: 40.0 (#163)

Long Context benchmarks
BenchmarkAmazon Nova ProGPT-4.1
LMArena Longer Query12551385
Fiction.LiveBench—63.9%

Writing & Preference GPT-4.1 leads

Amazon Nova Pro: 43.9 (#226), GPT-4.1: 57.6 (#125)

Writing & Preference benchmarks
BenchmarkAmazon Nova ProGPT-4.1
LMArena Text12591383
LMArena Creative Writing12121363
WildBench77.7%85.4%
LMArena Multi-Turn12461398
Short-Story Creative Writing60.5%—
EQ-Bench Creative Writing—1420
LiveBench Language37%—

Frequently asked questions

Is Amazon Nova Pro better than GPT-4.1?

GPT-4.1 is the stronger model overall, scoring 35.9 to 31.0 on the Noometry Index. Amazon Nova Pro costs 2.5× less per token, which makes it the better buy when GPT-4.1's lead doesn't matter for your workload.

Which is cheaper, Amazon Nova Pro or GPT-4.1?

Amazon Nova Pro is cheaper. It lists at $0.80 per million input tokens and $3.20 per million output tokens; GPT-4.1 lists at $2 and $8.

Is Amazon Nova Pro or GPT-4.1 better for coding?

They score almost the same on coding (35.1 vs 34.4); test both on your own repository before choosing.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 300K.

How many benchmarks do Amazon Nova Pro and GPT-4.1 share?

27 benchmarks have published results for both models. Amazon Nova Pro has 38 scored results on Noometry and GPT-4.1 has 52.

Related comparisons

Go deeper