Model comparison

Amazon Nova Micro vs Llama-3.3-70B-Instruct

Amazon Nova Micro and Llama-3.3-70B-Instruct score almost the same on the Noometry Index (30.4 vs 30.6), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Amazon Nova Micro scores higher in 3 categories and Llama-3.3-70B-Instruct in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama-3.3-70B-Instruct leads 71.1 to 56.3.
  • The biggest single-benchmark swing is LiveBench Instruction Following: 48% for Amazon Nova Micro and 82.7% for Llama-3.3-70B-Instruct.
  • Amazon Nova Micro is cheaper at $0.035 / $0.14 per million input/output tokens, against $0.10 / $0.32 for Llama-3.3-70B-Instruct.
  • Llama-3.3-70B-Instruct has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Micro and Llama-3.3-70B-Instruct specifications
Amazon Nova MicroLlama-3.3-70B-Instruct
ProviderAmazonMeta
Noometry Index30.430.6
Released2024-12-032024-12-06
WeightsProprietaryOpen
Context window128K128K
Max output10K4K
Input $ / M tokens$0.035$0.10
Output $ / M tokens$0.14$0.32
Results tracked3243

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Amazon Nova Micro: 30.5 (#295), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LiveBench Coding20.2%36.6%
LMArena Coding12181268
SciCode—26%
WeirdML—14.4%
BigCodeBench Instruct—46.9%
BigCodeBench Complete—57.5%

Agentic & Tool Use Llama-3.3-70B-Instruct leads

Amazon Nova Micro: 22.1 (#132), Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard22.3%31.9%
BALROG—23%

Reasoning Amazon Nova Micro leads

Amazon Nova Micro: 17.4 (#294), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LiveBench Reasoning25.1%50.8%
LMArena Hard Prompts11911257
LiveBench Data Analysis34%49.5%
LiveBench29.6%50.2%
SimpleBench—19.9%
CritPt—0%
DTBench—59.5%
LMCA—17.5%
Epoch Capabilities Index—127.33
ForecastBench—58.6

Math Amazon Nova Micro leads

Amazon Nova Micro: 26.9 (#254), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LiveBench Math34.5%42.2%
LMArena Math12061267
OTIS Mock AIME 2024-2025—5.1%
Omni-MATH21.4%—
MATH Level 5—41.6%

Knowledge Llama-3.3-70B-Instruct leads

Amazon Nova Micro: 29.6 (#237), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
Vectara Hallucination Rate5.5%4.1%
LMArena Expert11841225
MMLU70.8%86.3%
GPQA Diamond—47.4%
MMLU-Pro51.1%—
Confabulations—22.8%
GPQA (HELM)38.3%—

Multilingual Llama-3.3-70B-Instruct leads

Amazon Nova Micro: 36.5 (#239), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LMArena Non-English11861236
LMArena Chinese12091217
LMArena French12381281
LMArena German11921251
LMArena Japanese11541150
LMArena Korean11501143
LMArena Russian11851252
LMArena Spanish12251270

Instruction Following Llama-3.3-70B-Instruct leads

Amazon Nova Micro: 56.3 (#272), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LiveBench Instruction Following48%82.7%
LMArena Instruction Following11741242
IFEval76%—

Long Context Amazon Nova Micro leads

Amazon Nova Micro: 36.5 (#229), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LMArena Longer Query12051256
Fiction.LiveBench—33.3%

Writing & Preference Llama-3.3-70B-Instruct leads

Amazon Nova Micro: 39.5 (#247), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroLlama-3.3-70B-Instruct
LMArena Text12081274
LMArena Creative Writing11721250
LMArena Multi-Turn11781280
LiveBench Language15.8%39.2%
WildBench74.3%—

Frequently asked questions

Is Amazon Nova Micro better than Llama-3.3-70B-Instruct?

Amazon Nova Micro and Llama-3.3-70B-Instruct score almost the same on the Noometry Index (30.4 vs 30.6), so choose on price, context window or the category you care about most.

Which is cheaper, Amazon Nova Micro or Llama-3.3-70B-Instruct?

Amazon Nova Micro is cheaper. It lists at $0.035 per million input tokens and $0.14 per million output tokens; Llama-3.3-70B-Instruct lists at $0.10 and $0.32.

Is Amazon Nova Micro or Llama-3.3-70B-Instruct better for coding?

They score almost the same on coding (30.5 vs 31.0); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Amazon Nova Micro and Llama-3.3-70B-Instruct share?

27 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper