Model comparison

Amazon Nova Pro vs phi-3-medium 14B

Amazon Nova Pro is the stronger model overall, scoring 31.0 to 29.7 on the Noometry Index.

Last verified . 2 shared benchmarks.

Amazon Nova Pro Amazon

31.0

Rank #281 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. Amazon Nova Pro scores higher in 2 categories and phi-3-medium 14B in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Amazon Nova Pro leads 27.4 to 9.1.
  • phi-3-medium 14B has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Pro and phi-3-medium 14B specifications
Amazon Nova Prophi-3-medium 14B
ProviderAmazonMicrosoft
Noometry Index31.029.7
Released2024-12-032024-04-23
WeightsProprietaryOpen
Context window300K—
Max output10K—
Input $ / M tokens$0.80—
Output $ / M tokens$3.20—
Results tracked3813

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

Amazon Nova Pro: 35.1 (#229), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
BigCodeBench Instruct—37.6%
LiveBench Coding38.1%—
LMArena Coding1270—
BigCodeBench Complete—48.7%

Agentic & Tool Use Not comparable

Amazon Nova Pro: 16.7 (#147), phi-3-medium 14B: —

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
Berkeley Function Calling Leaderboard25%—
TheAgentCompany1.7%—

Reasoning Not comparable

Amazon Nova Pro: 20.0 (#243), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
Epoch Capabilities Index123.8121.23
LiveBench Reasoning32.6%—
LMArena Hard Prompts1246—
LiveBench Data Analysis48.3%—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
HellaSwag—82.4%
LiveBench43.5%—
WinoGrande—81.5%

Math Amazon Nova Pro leads

Amazon Nova Pro: 28.5 (#243), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
Omni-MATH24.2%—
LiveBench Math38%—
LMArena Math1252—
MATH Level 5—17.6%

Knowledge Amazon Nova Pro leads

Amazon Nova Pro: 27.4 (#250), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
MMLU82%78%
GPQA Diamond—27.6%
Humanity's Last Exam4.4%—
MMLU-Pro67.3%—
Confabulations30.1%—
Vectara Hallucination Rate5.1%—
GPQA (HELM)44.6%—
LMArena Expert1211—
ARC (AI2) Challenge—91.6%
OpenBookQA—87.4%
TriviaQA—73.9%

Multimodal Not comparable

Amazon Nova Pro: 25.0 (#126), phi-3-medium 14B: —

Multimodal benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
LMArena Vision980—

Multilingual Not comparable

Amazon Nova Pro: 39.7 (#223), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
LMArena Non-English1234—
LMArena Chinese1244—
LMArena French1271—
LMArena German1243—
LMArena Japanese1200—
LMArena Korean1203—
LMArena Russian1240—
LMArena Spanish1182—

Instruction Following Not comparable

Amazon Nova Pro: 64.9 (#226), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
LiveBench Instruction Following67.1%—
IFEval81.5%—
LMArena Instruction Following1235—

Long Context Not comparable

Amazon Nova Pro: 38.1 (#205), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
LMArena Longer Query1255—

Writing & Preference Not comparable

Amazon Nova Pro: 43.9 (#226), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkAmazon Nova Prophi-3-medium 14B
LMArena Text1259—
LMArena Creative Writing1212—
Short-Story Creative Writing60.5%—
WildBench77.7%—
LMArena Multi-Turn1246—
LiveBench Language37%—

Frequently asked questions

Is Amazon Nova Pro better than phi-3-medium 14B?

Amazon Nova Pro is the stronger model overall, scoring 31.0 to 29.7 on the Noometry Index.

Is Amazon Nova Pro or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 35.1 in the Noometry coding category.

How many benchmarks do Amazon Nova Pro and phi-3-medium 14B share?

2 benchmarks have published results for both models. Amazon Nova Pro has 38 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper