Model comparison

Amazon Nova Micro vs Qwen3-4B

Qwen3-4B is the stronger model overall, scoring 31.9 to 30.4 on the Noometry Index.

Last verified . 2 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Amazon Nova Micro scores higher in 0 categories and Qwen3-4B in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Qwen3-4B leads 27.6 to 22.1.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 22.3% for Amazon Nova Micro and 35.7% for Qwen3-4B.
  • Qwen3-4B has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Micro and Qwen3-4B specifications
Amazon Nova MicroQwen3-4B
ProviderAmazonAlibaba (Qwen)
Noometry Index30.431.9
Released2024-12-032025-04-29
WeightsProprietaryOpen
Context window128K—
Max output10K—
Input $ / M tokens$0.035—
Output $ / M tokens$0.14—
Results tracked326

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Amazon Nova Micro: 30.5 (#295), Qwen3-4B: —

Coding benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
LiveBench Coding20.2%—
LMArena Coding1218—

Agentic & Tool Use Qwen3-4B leads

Amazon Nova Micro: 22.1 (#132), Qwen3-4B: 27.6 (#100)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
Berkeley Function Calling Leaderboard22.3%35.7%

Reasoning Qwen3-4B leads

Amazon Nova Micro: 17.4 (#294), Qwen3-4B: 19.2 (#268)

Reasoning benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
Chess Puzzles—4%
LiveBench Reasoning25.1%—
LMArena Hard Prompts1191—
LiveBench Data Analysis34%—
LiveBench29.6%—

Math Qwen3-4B leads

Amazon Nova Micro: 26.9 (#254), Qwen3-4B: 29.7 (#240)

Math benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
MathArena Final-Answer Competitions—38.5%
OTIS Mock AIME 2024-2025—52.2%
Omni-MATH21.4%—
LiveBench Math34.5%—
LMArena Math1206—

Knowledge Qwen3-4B leads

Amazon Nova Micro: 29.6 (#237), Qwen3-4B: 33.0 (#208)

Knowledge benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
Vectara Hallucination Rate5.5%5.7%
GPQA Diamond—52.3%
MMLU-Pro51.1%—
GPQA (HELM)38.3%—
LMArena Expert1184—
MMLU70.8%—

Multilingual Not comparable

Amazon Nova Micro: 36.5 (#239), Qwen3-4B: —

Multilingual benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
LMArena Non-English1186—
LMArena Chinese1209—
LMArena French1238—
LMArena German1192—
LMArena Japanese1154—
LMArena Korean1150—
LMArena Russian1185—
LMArena Spanish1225—

Instruction Following Not comparable

Amazon Nova Micro: 56.3 (#272), Qwen3-4B: —

Instruction Following benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
LiveBench Instruction Following48%—
IFEval76%—
LMArena Instruction Following1174—

Long Context Not comparable

Amazon Nova Micro: 36.5 (#229), Qwen3-4B: —

Long Context benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
LMArena Longer Query1205—

Writing & Preference Not comparable

Amazon Nova Micro: 39.5 (#247), Qwen3-4B: —

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroQwen3-4B
LMArena Text1208—
LMArena Creative Writing1172—
WildBench74.3%—
LMArena Multi-Turn1178—
LiveBench Language15.8%—

Frequently asked questions

Is Amazon Nova Micro better than Qwen3-4B?

Qwen3-4B is the stronger model overall, scoring 31.9 to 30.4 on the Noometry Index.

How many benchmarks do Amazon Nova Micro and Qwen3-4B share?

2 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Qwen3-4B has 6.

Related comparisons

Go deeper