Model comparison

Amazon Nova Micro vs Claude 3.7 Sonnet

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 30.4 on the Noometry Index.

Last verified . 29 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Amazon Nova Micro scores higher in 0 categories and Claude 3.7 Sonnet in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Claude 3.7 Sonnet leads 72.9 to 56.3.
  • The biggest single-benchmark swing is LiveBench Reasoning: 25.1% for Amazon Nova Micro and 87.8% for Claude 3.7 Sonnet.

Side by side

Amazon Nova Micro and Claude 3.7 Sonnet specifications
Amazon Nova MicroClaude 3.7 Sonnet
ProviderAmazonAnthropic
Noometry Index30.439.5
Released2024-12-032025-02-24
WeightsProprietaryProprietary
Context window128K—
Max output10K—
Input $ / M tokens$0.035—
Output $ / M tokens$0.14—
Results tracked3258

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Amazon Nova Micro: 30.5 (#295), Claude 3.7 Sonnet: 40.6 (#136)

Coding benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LiveBench Coding20.2%74.5%
LMArena Coding12181361
SWE-bench Verified—61%
SWE-bench Verified (bash only)—52.8%
Aider Polyglot—64.9%
GSO—3.8%
CadEval—54%

Agentic & Tool Use Claude 3.7 Sonnet leads

Amazon Nova Micro: 22.1 (#132), Claude 3.7 Sonnet: 34.1 (#50)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
Berkeley Function Calling Leaderboard22.3%—
TheAgentCompany—30.9%
Cybench—20%
DeepResearch Bench—43.6%
OSWorld—35.8%
METR Time Horizons—60%

Reasoning Claude 3.7 Sonnet leads

Amazon Nova Micro: 17.4 (#294), Claude 3.7 Sonnet: 18.6 (#277)

Reasoning benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LiveBench Reasoning25.1%87.8%
LMArena Hard Prompts11911333
LiveBench Data Analysis34%74%
LiveBench29.6%76.1%
ARC-AGI-2—0.9%
SimpleBench—46.4%
ARC-AGI-1—28.6%
EnigmaEval—4.2%
Epoch Capabilities Index—141.16
ForecastBench—61.8

Math Claude 3.7 Sonnet leads

Amazon Nova Micro: 26.9 (#254), Claude 3.7 Sonnet: 37.5 (#153)

Math benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
Omni-MATH21.4%33%
LiveBench Math34.5%79%
LMArena Math12061337
OTIS Mock AIME 2024-2025—57.8%
MATH Level 5—91.2%
FrontierMath (Feb 2025 set)—4.1%

Knowledge Claude 3.7 Sonnet leads

Amazon Nova Micro: 29.6 (#237), Claude 3.7 Sonnet: 39.8 (#130)

Knowledge benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
MMLU-Pro51.1%78.4%
GPQA (HELM)38.3%60.8%
LMArena Expert11841321
GPQA Diamond—79.7%
Humanity's Last Exam—8%
Confabulations—14.7%
Vectara Hallucination Rate5.5%—
MMLU70.8%—

Multimodal Not comparable

Amazon Nova Micro: —, Claude 3.7 Sonnet: 33.7 (#95)

Multimodal benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LMArena Vision—1169
GeoBench—68%
VPCT—39%
SpatialViz-Bench—33.9%

Multilingual Claude 3.7 Sonnet leads

Amazon Nova Micro: 36.5 (#239), Claude 3.7 Sonnet: 44.1 (#179)

Multilingual benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LMArena Non-English11861296
LMArena Chinese12091299
LMArena French12381303
LMArena German11921301
LMArena Japanese11541267
LMArena Korean11501249
LMArena Russian11851311
LMArena Spanish12251298

Instruction Following Claude 3.7 Sonnet leads

Amazon Nova Micro: 56.3 (#272), Claude 3.7 Sonnet: 72.9 (#125)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LiveBench Instruction Following48%81.3%
IFEval76%83.4%
LMArena Instruction Following11741352

Long Context Claude 3.7 Sonnet leads

Amazon Nova Micro: 36.5 (#229), Claude 3.7 Sonnet: 50.3 (#10)

Long Context benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LMArena Longer Query12051373
Fiction.LiveBench—83.3%

Writing & Preference Claude 3.7 Sonnet leads

Amazon Nova Micro: 39.5 (#247), Claude 3.7 Sonnet: 54.4 (#150)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroClaude 3.7 Sonnet
LMArena Text12081314
LMArena Creative Writing11721332
WildBench74.3%81.4%
LMArena Multi-Turn11781339
LiveBench Language15.8%59.9%
Short-Story Creative Writing—81.1%
EQ-Bench Creative Writing—1412

Frequently asked questions

Is Amazon Nova Micro better than Claude 3.7 Sonnet?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 30.4 on the Noometry Index.

Is Amazon Nova Micro or Claude 3.7 Sonnet better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 30.5 in the Noometry coding category.

How many benchmarks do Amazon Nova Micro and Claude 3.7 Sonnet share?

29 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Claude 3.7 Sonnet has 58.

Related comparisons

Go deeper