Model comparison

Amazon Nova Pro vs Grok-2 (Dec 2024)

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 31.0 on the Noometry Index.

Last verified . 27 shared benchmarks.

Amazon Nova Pro Amazon

31.0

Rank #281 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Amazon Nova Pro scores higher in 3 categories and Grok-2 (Dec 2024) in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Amazon Nova Pro leads 28.5 to 20.8.
  • The biggest single-benchmark swing is LiveBench Reasoning: 32.6% for Amazon Nova Pro and 54.8% for Grok-2 (Dec 2024).

Side by side

Amazon Nova Pro and Grok-2 (Dec 2024) specifications
Amazon Nova ProGrok-2 (Dec 2024)
ProviderAmazonxAI
Noometry Index31.033.7
Released2024-12-032024-08-13
WeightsProprietaryProprietary
Context window300K—
Max output10K—
Input $ / M tokens$0.80—
Output $ / M tokens$3.20—
Results tracked3834

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Amazon Nova Pro leads

Amazon Nova Pro: 35.1 (#229), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LiveBench Coding38.1%46.4%
LMArena Coding12701287
WeirdML—22.2%

Agentic & Tool Use Not comparable

Amazon Nova Pro: 16.7 (#147), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
Berkeley Function Calling Leaderboard25%—
TheAgentCompany1.7%—

Reasoning Amazon Nova Pro leads

Amazon Nova Pro: 20.0 (#243), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LiveBench Reasoning32.6%54.8%
LMArena Hard Prompts12461272
LiveBench Data Analysis48.3%54.5%
Epoch Capabilities Index123.8130.48
LiveBench43.5%54.3%
SimpleBench—22.7%
DTBench—65.2%

Math Amazon Nova Pro leads

Amazon Nova Pro: 28.5 (#243), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LiveBench Math38%54.9%
LMArena Math12521283
OTIS Mock AIME 2024-2025—11.5%
Omni-MATH24.2%—
MATH Level 5—63.5%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Grok-2 (Dec 2024) leads

Amazon Nova Pro: 27.4 (#250), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
Confabulations30.1%20.1%
LMArena Expert12111254
GPQA Diamond—53.8%
Humanity's Last Exam4.4%—
MMLU-Pro67.3%—
Vectara Hallucination Rate5.1%—
GPQA (HELM)44.6%—
MMLU82%—

Multimodal Not comparable

Amazon Nova Pro: 25.0 (#126), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LMArena Vision980—

Multilingual Grok-2 (Dec 2024) leads

Amazon Nova Pro: 39.7 (#223), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LMArena Non-English12341282
LMArena Chinese12441289
LMArena French12711318
LMArena German12431287
LMArena Japanese12001244
LMArena Korean12031237
LMArena Russian12401286
LMArena Spanish11821281

Instruction Following Grok-2 (Dec 2024) leads

Amazon Nova Pro: 64.9 (#226), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LiveBench Instruction Following67.1%69.6%
LMArena Instruction Following12351270
IFEval81.5%—

Long Context Too close to call

Amazon Nova Pro: 38.1 (#205), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LMArena Longer Query12551276

Writing & Preference Grok-2 (Dec 2024) leads

Amazon Nova Pro: 43.9 (#226), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkAmazon Nova ProGrok-2 (Dec 2024)
LMArena Text12591305
LMArena Creative Writing12121284
Short-Story Creative Writing60.5%63.6%
LMArena Multi-Turn12461290
LiveBench Language37%45.6%
WildBench77.7%—

Frequently asked questions

Is Amazon Nova Pro better than Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 31.0 on the Noometry Index.

Is Amazon Nova Pro or Grok-2 (Dec 2024) better for coding?

Amazon Nova Pro scores higher on coding benchmarks: 35.1 versus 33.3 in the Noometry coding category.

How many benchmarks do Amazon Nova Pro and Grok-2 (Dec 2024) share?

27 benchmarks have published results for both models. Amazon Nova Pro has 38 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper