Model comparison

Amazon Nova Micro vs Grok 3

Grok 3 is the stronger model overall, scoring 39.9 to 30.4 on the Noometry Index.

Last verified . 23 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Grok 3 xAI

39.9

Rank #157 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Amazon Nova Micro scores higher in 1 category and Grok 3 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Grok 3 leads 75.0 to 56.3.
  • The biggest single-benchmark swing is MMLU-Pro: 51.1% for Amazon Nova Micro and 78.8% for Grok 3.

Side by side

Amazon Nova Micro and Grok 3 specifications
Amazon Nova MicroGrok 3
ProviderAmazonxAI
Noometry Index30.439.9
Released2024-12-032025-04-09
WeightsProprietaryProprietary
Context window128K—
Max output10K—
Input $ / M tokens$0.035—
Output $ / M tokens$0.14—
Results tracked3240

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 3 leads

Amazon Nova Micro: 30.5 (#295), Grok 3: 41.9 (#115)

Coding benchmarks
BenchmarkAmazon Nova MicroGrok 3
LMArena Coding12181432
Aider Polyglot—53.3%
WeirdML—37.2%
LiveBench Coding20.2%—

Agentic & Tool Use Grok 3 leads

Amazon Nova Micro: 22.1 (#132), Grok 3: 30.5 (#76)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroGrok 3
Berkeley Function Calling Leaderboard22.3%—
BALROG—29.5%

Reasoning Amazon Nova Micro leads

Amazon Nova Micro: 17.4 (#294), Grok 3: 13.7 (#333)

Reasoning benchmarks
BenchmarkAmazon Nova MicroGrok 3
LMArena Hard Prompts11911434
ARC-AGI-2—0%
SimpleBench—36.1%
Kagi LLM Benchmark—61.3%
ARC-AGI-1—5.5%
LiveBench Reasoning25.1%—
LiveBench Data Analysis34%—
Epoch Capabilities Index—138.33
LiveBench29.6%—

Math Grok 3 leads

Amazon Nova Micro: 26.9 (#254), Grok 3: 38.0 (#145)

Math benchmarks
BenchmarkAmazon Nova MicroGrok 3
Omni-MATH21.4%46.4%
LMArena Math12061391
OTIS Mock AIME 2024-2025—55.6%
LiveBench Math34.5%—
MATH Level 5—88.7%
FrontierMath (Feb 2025 set)—3.8%
FrontierMath Tier 4 (v1)—0%

Knowledge Grok 3 leads

Amazon Nova Micro: 29.6 (#237), Grok 3: 46.2 (#82)

Knowledge benchmarks
BenchmarkAmazon Nova MicroGrok 3
MMLU-Pro51.1%78.8%
Vectara Hallucination Rate5.5%5.8%
GPQA (HELM)38.3%65%
LMArena Expert11841421
GPQA Diamond—75.8%
Confabulations—14.2%
MMLU70.8%—

Multilingual Grok 3 leads

Amazon Nova Micro: 36.5 (#239), Grok 3: 52.3 (#87)

Multilingual benchmarks
BenchmarkAmazon Nova MicroGrok 3
LMArena Non-English11861410
LMArena Chinese12091448
LMArena French12381460
LMArena German11921431
LMArena Japanese11541387
LMArena Korean11501373
LMArena Russian11851416
LMArena Spanish12251417

Instruction Following Grok 3 leads

Amazon Nova Micro: 56.3 (#272), Grok 3: 75.0 (#73)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroGrok 3
IFEval76%88.4%
LMArena Instruction Following11741409
LiveBench Instruction Following48%—

Long Context Grok 3 leads

Amazon Nova Micro: 36.5 (#229), Grok 3: 38.7 (#192)

Long Context benchmarks
BenchmarkAmazon Nova MicroGrok 3
LMArena Longer Query12051439
Fiction.LiveBench—58.3%

Writing & Preference Grok 3 leads

Amazon Nova Micro: 39.5 (#247), Grok 3: 55.8 (#141)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroGrok 3
LMArena Text12081426
LMArena Creative Writing11721414
WildBench74.3%84.9%
LMArena Multi-Turn11781425
Short-Story Creative Writing—76.4%
EQ-Bench Creative Writing—1186
LiveBench Language15.8%—

Frequently asked questions

Is Amazon Nova Micro better than Grok 3?

Grok 3 is the stronger model overall, scoring 39.9 to 30.4 on the Noometry Index.

Is Amazon Nova Micro or Grok 3 better for coding?

Grok 3 scores higher on coding benchmarks: 41.9 versus 30.5 in the Noometry coding category.

How many benchmarks do Amazon Nova Micro and Grok 3 share?

23 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Grok 3 has 40.

Related comparisons

Go deeper