Model comparison

Amazon Nova Experimental Chat 26 02 10 vs GPT-5

GPT-5 is the stronger model overall, scoring 50.9 to 44.5 on the Noometry Index.

Last verified . 12 shared benchmarks.

GPT-5 OpenAI

50.9

Rank #45 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Amazon Nova Experimental Chat 26 02 10 scores higher in 2 categories and GPT-5 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where GPT-5 leads 69.5 to 44.0.

Side by side

Amazon Nova Experimental Chat 26 02 10 and GPT-5 specifications
Amazon Nova Experimental Chat 26 02 10GPT-5
ProviderAmazonOpenAI
Noometry Index44.550.9
Released—2025-08-07
WeightsProprietaryProprietary
Context window—400K
Max output—128K
Input $ / M tokens—$1.25
Output $ / M tokens—$10
Results tracked1269

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 leads

Amazon Nova Experimental Chat 26 02 10: 43.9 (#82), GPT-5: 50.3 (#47)

Coding benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Coding14831436
SWE-bench Verified—73.6%
SWE-bench Verified (bash only)—65%
Aider Polyglot—88%
LMArena WebDev—1418
SciCode—42.9%
GSO—6.9%
WeirdML—60.7%
ALE-Bench—1,162
AlgoTune—1.67

Agentic & Tool Use Not comparable

Amazon Nova Experimental Chat 26 02 10: —, GPT-5: 33.1 (#56)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
Terminal-Bench—49.6%
GDPval—34.8%
Remote Labor Index—1.7%
DeepResearch Bench—49.6%
BALROG—32.8%
LMArena Search—1133
METR Time Horizons—69.6%

Reasoning GPT-5 leads

Amazon Nova Experimental Chat 26 02 10: 30.1 (#86), GPT-5: 38.3 (#64)

Reasoning benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Hard Prompts14581416
ARC-AGI-2—9.9%
SimpleBench—56.7%
Kagi LLM Benchmark—72.7%
ARC-AGI-1—65.7%
CritPt—12.6%
Chess Puzzles—37%
EnigmaEval—10.5%
EBR-Bench—12.7%
Mystery Game Puzzles—23%
DTBench—90.7%
LMCA—40%
Epoch Capabilities Index—150
ForecastBench—61.4

Math GPT-5 leads

Amazon Nova Experimental Chat 26 02 10: 39.3 (#109), GPT-5: 55.0 (#44)

Math benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Math14381407
FrontierMath (Tiers 1-3)—55.4%
FrontierMath Tier 4—22%
OTIS Mock AIME 2024-2025—91.4%
ProofBench—18%
Omni-MATH—64.7%
MATH Level 5—98.1%
FrontierMath (Feb 2025 set)—32.4%
FrontierMath Tier 4 (v1)—12.5%

Knowledge GPT-5 leads

Amazon Nova Experimental Chat 26 02 10: 42.3 (#97), GPT-5: 56.6 (#43)

Knowledge benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Expert15061419
GPQA Diamond—86.2%
Humanity's Last Exam—25.3%
SimpleQA Verified—50.1%
MMLU-Pro—86.3%
Confabulations—10.3%
Vectara Hallucination Rate—14.7%
GPQA (HELM)—79.2%

Multimodal Not comparable

Amazon Nova Experimental Chat 26 02 10: —, GPT-5: 46.8 (#13)

Multimodal benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Vision—1232
GeoBench—81%
VPCT—66%

Multilingual Amazon Nova Experimental Chat 26 02 10 leads

Amazon Nova Experimental Chat 26 02 10: 54.0 (#50), GPT-5: 51.4 (#110)

Multilingual benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Non-English14341397
LMArena Chinese14631422
LMArena Russian14231406
LMArena French—1410
LMArena German—1416
LMArena Japanese—1409
LMArena Korean—1360
LMArena Spanish—1399

Instruction Following Amazon Nova Experimental Chat 26 02 10 leads

Amazon Nova Experimental Chat 26 02 10: 75.2 (#69), GPT-5: 73.8 (#113)

Instruction Following benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Instruction Following14271388
IFEval—87.5%

Long Context GPT-5 leads

Amazon Nova Experimental Chat 26 02 10: 44.0 (#77), GPT-5: 69.5 (#2)

Long Context benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Longer Query14411399
Fiction.LiveBench—97.2%

Writing & Preference GPT-5 leads

Amazon Nova Experimental Chat 26 02 10: 61.8 (#84), GPT-5: 63.4 (#65)

Writing & Preference benchmarks
BenchmarkAmazon Nova Experimental Chat 26 02 10GPT-5
LMArena Text14481406
LMArena Creative Writing13681365
LMArena Multi-Turn14421426
Short-Story Creative Writing—86%
EQ-Bench Creative Writing—1627
WildBench—85.7%

Frequently asked questions

Is Amazon Nova Experimental Chat 26 02 10 better than GPT-5?

GPT-5 is the stronger model overall, scoring 50.9 to 44.5 on the Noometry Index.

Is Amazon Nova Experimental Chat 26 02 10 or GPT-5 better for coding?

GPT-5 scores higher on coding benchmarks: 50.3 versus 43.9 in the Noometry coding category.

How many benchmarks do Amazon Nova Experimental Chat 26 02 10 and GPT-5 share?

12 benchmarks have published results for both models. Amazon Nova Experimental Chat 26 02 10 has 12 scored results on Noometry and GPT-5 has 69.

Related comparisons

Go deeper