Model comparison

Amazon Nova Micro vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.4 on the Noometry Index.

Last verified . 9 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Amazon Nova Micro scores higher in 5 categories and Llama2 70b Steerlm Chat in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Amazon Nova Micro leads 39.5 to 31.6.
  • Llama2 70b Steerlm Chat has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Micro and Llama2 70b Steerlm Chat specifications
Amazon Nova MicroLlama2 70b Steerlm Chat
ProviderAmazonNVIDIA
Noometry Index30.431.8
Released2024-12-03—
WeightsProprietaryOpen
Context window128K—
Max output10K—
Input $ / M tokens$0.035—
Output $ / M tokens$0.14—
Results tracked329

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Amazon Nova Micro: 30.5 (#295), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Coding12181025
LiveBench Coding20.2%—

Agentic & Tool Use Not comparable

Amazon Nova Micro: 22.1 (#132), Llama2 70b Steerlm Chat: —

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
Berkeley Function Calling Leaderboard22.3%—

Reasoning Llama2 70b Steerlm Chat leads

Amazon Nova Micro: 17.4 (#294), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Hard Prompts11911047
LiveBench Reasoning25.1%—
LiveBench Data Analysis34%—
LiveBench29.6%—

Math Llama2 70b Steerlm Chat leads

Amazon Nova Micro: 26.9 (#254), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Math12061072
Omni-MATH21.4%—
LiveBench Math34.5%—

Knowledge Not comparable

Amazon Nova Micro: 29.6 (#237), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
MMLU-Pro51.1%—
Vectara Hallucination Rate5.5%—
GPQA (HELM)38.3%—
LMArena Expert1184—
MMLU70.8%—

Multilingual Amazon Nova Micro leads

Amazon Nova Micro: 36.5 (#239), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Non-English11861063
LMArena Chinese1209—
LMArena French1238—
LMArena German1192—
LMArena Japanese1154—
LMArena Korean1150—
LMArena Russian1185—
LMArena Spanish1225—

Instruction Following Amazon Nova Micro leads

Amazon Nova Micro: 56.3 (#272), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Instruction Following11741060
LiveBench Instruction Following48%—
IFEval76%—

Long Context Amazon Nova Micro leads

Amazon Nova Micro: 36.5 (#229), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Longer Query1205998

Writing & Preference Amazon Nova Micro leads

Amazon Nova Micro: 39.5 (#247), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroLlama2 70b Steerlm Chat
LMArena Text12081098
LMArena Creative Writing11721091
LMArena Multi-Turn11781058
WildBench74.3%—
LiveBench Language15.8%—

Frequently asked questions

Is Amazon Nova Micro better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.4 on the Noometry Index.

Is Amazon Nova Micro or Llama2 70b Steerlm Chat better for coding?

They score almost the same on coding (30.5 vs 29.9); test both on your own repository before choosing.

How many benchmarks do Amazon Nova Micro and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper