Model comparison

Amazon Nova Micro vs Qwen2.5-Coder-32B

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 30.4 on the Noometry Index. Amazon Nova Micro costs 12× less per token, which makes it the better buy when Qwen2.5-Coder-32B's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Amazon Nova Micro scores higher in 1 category and Qwen2.5-Coder-32B in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Amazon Nova Micro leads 30.5 to 22.6.
  • The biggest single-benchmark swing is LiveBench Coding: 20.2% for Amazon Nova Micro and 56.9% for Qwen2.5-Coder-32B.
  • Amazon Nova Micro is cheaper at $0.035 / $0.14 per million input/output tokens, against $0.66 / $1 for Qwen2.5-Coder-32B.
  • Amazon Nova Micro accepts more context: 128K tokens versus 33K.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Micro and Qwen2.5-Coder-32B specifications
Amazon Nova MicroQwen2.5-Coder-32B
ProviderAmazonAlibaba (Qwen)
Noometry Index30.433.4
Released2024-12-032024-09-18
WeightsProprietaryOpen
Context window128K33K
Max output10K29K
Input $ / M tokens$0.035$0.66
Output $ / M tokens$0.14$1
Results tracked3231

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Amazon Nova Micro leads

Amazon Nova Micro: 30.5 (#295), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LiveBench Coding20.2%56.9%
LMArena Coding12181276
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
BigCodeBench Instruct—49%
BigCodeBench Complete—58%
HumanEval+—87.2%
MBPP+—77%

Agentic & Tool Use Not comparable

Amazon Nova Micro: 22.1 (#132), Qwen2.5-Coder-32B: —

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
Berkeley Function Calling Leaderboard22.3%—

Reasoning Qwen2.5-Coder-32B leads

Amazon Nova Micro: 17.4 (#294), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LiveBench Reasoning25.1%42.1%
LMArena Hard Prompts11911251
LiveBench Data Analysis34%49.9%
LiveBench29.6%46.2%
Epoch Capabilities Index—119.49
HellaSwag—83%
WinoGrande—80.8%

Math Qwen2.5-Coder-32B leads

Amazon Nova Micro: 26.9 (#254), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LiveBench Math34.5%46.6%
LMArena Math12061251
Omni-MATH21.4%—
GSM8K—93%

Knowledge Qwen2.5-Coder-32B leads

Amazon Nova Micro: 29.6 (#237), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LMArena Expert11841221
MMLU70.8%79.1%
MMLU-Pro51.1%—
Vectara Hallucination Rate5.5%—
GPQA (HELM)38.3%—
ARC (AI2) Challenge—70.5%

Multilingual Qwen2.5-Coder-32B leads

Amazon Nova Micro: 36.5 (#239), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LMArena Non-English11861205
LMArena Chinese12091222
LMArena Russian11851228
LMArena French1238—
LMArena German1192—
LMArena Japanese1154—
LMArena Korean1150—
LMArena Spanish1225—

Instruction Following Qwen2.5-Coder-32B leads

Amazon Nova Micro: 56.3 (#272), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LiveBench Instruction Following48%58.7%
LMArena Instruction Following11741223
IFEval76%—

Long Context Qwen2.5-Coder-32B leads

Amazon Nova Micro: 36.5 (#229), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LMArena Longer Query12051251

Writing & Preference Qwen2.5-Coder-32B leads

Amazon Nova Micro: 39.5 (#247), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroQwen2.5-Coder-32B
LMArena Text12081230
LMArena Creative Writing11721174
LMArena Multi-Turn11781222
LiveBench Language15.8%23.3%
WildBench74.3%—

Frequently asked questions

Is Amazon Nova Micro better than Qwen2.5-Coder-32B?

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 30.4 on the Noometry Index. Amazon Nova Micro costs 12× less per token, which makes it the better buy when Qwen2.5-Coder-32B's lead doesn't matter for your workload.

Which is cheaper, Amazon Nova Micro or Qwen2.5-Coder-32B?

Amazon Nova Micro is cheaper. It lists at $0.035 per million input tokens and $0.14 per million output tokens; Qwen2.5-Coder-32B lists at $0.66 and $1.

Is Amazon Nova Micro or Qwen2.5-Coder-32B better for coding?

Amazon Nova Micro scores higher on coding benchmarks: 30.5 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

Amazon Nova Micro does, with 128K tokens against 33K.

How many benchmarks do Amazon Nova Micro and Qwen2.5-Coder-32B share?

20 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper