Model comparison

Amazon Nova Micro vs GPT-4o

Amazon Nova Micro is the stronger model overall, scoring 30.4 to 28.6 on the Noometry Index.

Last verified . 31 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

GPT-4o OpenAI

28.6

Rank #324 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Amazon Nova Micro scores higher in 5 categories and GPT-4o in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Amazon Nova Micro leads 26.9 to 10.6.
  • The biggest single-benchmark swing is LiveBench Language: 15.8% for Amazon Nova Micro and 47.6% for GPT-4o.
  • Amazon Nova Micro is cheaper at $0.035 / $0.14 per million input/output tokens, against $2.50 / $10 for GPT-4o.

Side by side

Amazon Nova Micro and GPT-4o specifications
Amazon Nova MicroGPT-4o
ProviderAmazonOpenAI
Noometry Index30.428.6
Released2024-12-032024-05-13
WeightsProprietaryProprietary
Context window128K128K
Max output10K16K
Input $ / M tokens$0.035$2.50
Output $ / M tokens$0.14$10
Results tracked3272

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Amazon Nova Micro leads

Amazon Nova Micro: 30.5 (#295), GPT-4o: 24.8 (#328)

Coding benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LiveBench Coding20.2%51.4%
LMArena Coding12181297
SWE-bench Verified—31%
SWE-bench Verified (bash only)—21.6%
Aider Polyglot—45.3%
GSO—0%
WeirdML—25.1%
BigCodeBench Instruct—51.1%
BigCodeBench Complete—61.1%
CadEval—26%
HumanEval+—87.2%
MBPP+—72.2%

Agentic & Tool Use Amazon Nova Micro leads

Amazon Nova Micro: 22.1 (#132), GPT-4o: 21.0 (#141)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroGPT-4o
Berkeley Function Calling Leaderboard22.3%—
GDPval—9.9%
TheAgentCompany—8.6%
Cybench—12.5%
BALROG—32.3%
LMArena Search—1006
METR Time Horizons—40.8%

Reasoning Amazon Nova Micro leads

Amazon Nova Micro: 17.4 (#294), GPT-4o: 9.4 (#343)

Reasoning benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LiveBench Reasoning25.1%55.8%
LMArena Hard Prompts11911281
LiveBench Data Analysis34%60.9%
LiveBench29.6%55.3%
ARC-AGI-2—0%
SimpleBench—17.8%
ARC-AGI-1—4.5%
CritPt—0%
Chess Puzzles—13%
EnigmaEval—0.8%
DTBench—64.5%
LMCA—16.6%
Epoch Capabilities Index—128.97
ForecastBench—57.7

Math Amazon Nova Micro leads

Amazon Nova Micro: 26.9 (#254), GPT-4o: 10.6 (#312)

Math benchmarks
BenchmarkAmazon Nova MicroGPT-4o
Omni-MATH21.4%29.3%
LiveBench Math34.5%49.5%
LMArena Math12061285
FrontierMath (Tiers 1-3)—0.4%
OTIS Mock AIME 2024-2025—6.4%
MATH Level 5—53.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Too close to call

Amazon Nova Micro: 29.6 (#237), GPT-4o: 28.8 (#242)

Knowledge benchmarks
BenchmarkAmazon Nova MicroGPT-4o
MMLU-Pro51.1%71.3%
Vectara Hallucination Rate5.5%9.6%
GPQA (HELM)38.3%52%
LMArena Expert11841250
MMLU70.8%88.1%
GPQA Diamond—49.2%
Humanity's Last Exam—2.7%
SimpleQA Verified—26%
Confabulations—15.3%

Multimodal Not comparable

Amazon Nova Micro: —, GPT-4o: 34.5 (#91)

Multimodal benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LMArena Vision—1137
Video-MME—71.9%
GeoBench—71%
VPCT—40%
ScienceQA—88.5%

Multilingual GPT-4o leads

Amazon Nova Micro: 36.5 (#239), GPT-4o: 43.2 (#186)

Multilingual benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LMArena Non-English11861283
LMArena Chinese12091277
LMArena French12381304
LMArena German11921282
LMArena Japanese11541257
LMArena Korean11501234
LMArena Russian11851286
LMArena Spanish12251292

Instruction Following GPT-4o leads

Amazon Nova Micro: 56.3 (#272), GPT-4o: 66.6 (#207)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LiveBench Instruction Following48%68.6%
IFEval76%81.7%
LMArena Instruction Following11741278

Long Context GPT-4o leads

Amazon Nova Micro: 36.5 (#229), GPT-4o: 39.4 (#179)

Long Context benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LMArena Longer Query12051289
Fiction.LiveBench—66.7%

Writing & Preference GPT-4o leads

Amazon Nova Micro: 39.5 (#247), GPT-4o: 52.6 (#166)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroGPT-4o
LMArena Text12081300
LMArena Creative Writing11721292
WildBench74.3%82.8%
LMArena Multi-Turn11781302
LiveBench Language15.8%47.6%
Short-Story Creative Writing—81.8%

Frequently asked questions

Is Amazon Nova Micro better than GPT-4o?

Amazon Nova Micro is the stronger model overall, scoring 30.4 to 28.6 on the Noometry Index.

Which is cheaper, Amazon Nova Micro or GPT-4o?

Amazon Nova Micro is cheaper. It lists at $0.035 per million input tokens and $0.14 per million output tokens; GPT-4o lists at $2.50 and $10.

Is Amazon Nova Micro or GPT-4o better for coding?

Amazon Nova Micro scores higher on coding benchmarks: 30.5 versus 24.8 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Amazon Nova Micro and GPT-4o share?

31 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and GPT-4o has 72.

Related comparisons

Go deeper