Model comparison

Amazon Nova Lite vs GPT-5.1-Codex

GPT-5.1-Codex is the stronger model overall, scoring 38.6 to 31.9 on the Noometry Index. Amazon Nova Lite costs 33× less per token, which makes it the better buy when GPT-5.1-Codex's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Amazon Nova Lite Amazon

31.9

Rank #265 Confirmed

GPT-5.1-Codex OpenAI

38.6

Rank #186 Reported

Summary

  • They share 1 benchmark with published results for both. Amazon Nova Lite scores higher in 0 categories and GPT-5.1-Codex in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GPT-5.1-Codex leads 41.9 to 32.5.
  • Amazon Nova Lite is cheaper at $0.06 / $0.24 per million input/output tokens, against $1.25 / $10 for GPT-5.1-Codex.
  • GPT-5.1-Codex accepts more context: 400K tokens versus 300K.

Side by side

Amazon Nova Lite and GPT-5.1-Codex specifications
Amazon Nova LiteGPT-5.1-Codex
ProviderAmazonOpenAI
Noometry Index31.938.6
Released2024-12-032025-11-12
WeightsProprietaryProprietary
Context window300K400K
Max output10K128K
Input $ / M tokens$0.06$1.25
Output $ / M tokens$0.24$10
Results tracked346

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1-Codex leads

Amazon Nova Lite: 32.5 (#270), GPT-5.1-Codex: 41.9 (#116)

Coding benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
ALE-Bench236.251,245
SWE-bench Verified (bash only)—66%
LMArena WebDev—1337
LiveBench Coding27.5%—
LMArena Coding1239—

Agentic & Tool Use Not comparable

Amazon Nova Lite: —, GPT-5.1-Codex: 38.0 (#33)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
Terminal-Bench—60.4%
METR Time Horizons—70.8%

Reasoning Not comparable

Amazon Nova Lite: 19.3 (#260), GPT-5.1-Codex: —

Reasoning benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
LiveBench Reasoning36.7%—
LMArena Hard Prompts1220—
LiveBench Data Analysis37.2%—
LiveBench36.4%—

Math GPT-5.1-Codex leads

Amazon Nova Lite: 27.8 (#247), GPT-5.1-Codex: 30.3 (#235)

Math benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
ProofBench—9%
Omni-MATH23.3%—
LiveBench Math36.7%—
LMArena Math1227—

Knowledge Not comparable

Amazon Nova Lite: 26.3 (#259), GPT-5.1-Codex: —

Knowledge benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
Humanity's Last Exam3.6%—
MMLU-Pro60%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)39.7%—
LMArena Expert1201—
MMLU77%—

Multimodal Not comparable

Amazon Nova Lite: 25.5 (#123), GPT-5.1-Codex: —

Multimodal benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
LMArena Vision990—

Multilingual Not comparable

Amazon Nova Lite: 38.0 (#232), GPT-5.1-Codex: —

Multilingual benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
LMArena Non-English1208—
LMArena Chinese1225—
LMArena French1237—
LMArena German1229—
LMArena Japanese1153—
LMArena Korean1153—
LMArena Russian1216—
LMArena Spanish1226—

Instruction Following Not comparable

Amazon Nova Lite: 59.3 (#255), GPT-5.1-Codex: —

Instruction Following benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
LiveBench Instruction Following54.1%—
IFEval77.6%—
LMArena Instruction Following1205—

Long Context Not comparable

Amazon Nova Lite: 37.4 (#217), GPT-5.1-Codex: —

Long Context benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
LMArena Longer Query1234—

Writing & Preference Not comparable

Amazon Nova Lite: 42.1 (#237), GPT-5.1-Codex: —

Writing & Preference benchmarks
BenchmarkAmazon Nova LiteGPT-5.1-Codex
LMArena Text1229—
LMArena Creative Writing1197—
WildBench75%—
LMArena Multi-Turn1199—
LiveBench Language25.9%—

Frequently asked questions

Is Amazon Nova Lite better than GPT-5.1-Codex?

GPT-5.1-Codex is the stronger model overall, scoring 38.6 to 31.9 on the Noometry Index. Amazon Nova Lite costs 33× less per token, which makes it the better buy when GPT-5.1-Codex's lead doesn't matter for your workload.

Which is cheaper, Amazon Nova Lite or GPT-5.1-Codex?

Amazon Nova Lite is cheaper. It lists at $0.06 per million input tokens and $0.24 per million output tokens; GPT-5.1-Codex lists at $1.25 and $10.

Is Amazon Nova Lite or GPT-5.1-Codex better for coding?

GPT-5.1-Codex scores higher on coding benchmarks: 41.9 versus 32.5 in the Noometry coding category.

Which has the bigger context window?

GPT-5.1-Codex does, with 400K tokens against 300K.

How many benchmarks do Amazon Nova Lite and GPT-5.1-Codex share?

1 benchmark has published results for both models. Amazon Nova Lite has 34 scored results on Noometry and GPT-5.1-Codex has 6.

Related comparisons

Go deeper