Model comparison

Nova 2 Lite vs phi-3-medium 14B

Nova 2 Lite is the stronger model overall, scoring 39.7 to 29.7 on the Noometry Index.

Last verified . 0 shared benchmarks.

Nova 2 Lite Amazon

39.7

Rank #161 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • The widest gap is in knowledge, where Nova 2 Lite leads 43.0 to 9.1.
  • phi-3-medium 14B has downloadable open weights; the other is API-only.

Side by side

Nova 2 Lite and phi-3-medium 14B specifications
Nova 2 Litephi-3-medium 14B
ProviderAmazonMicrosoft
Noometry Index39.729.7
Released2025-12-012024-04-23
WeightsProprietaryOpen
Context window1M—
Max output64K—
Input $ / M tokens$0.30—
Output $ / M tokens$2.50—
Results tracked1913

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nova 2 Lite leads

Nova 2 Lite: 40.7 (#134), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
BigCodeBench Instruct—37.6%
LMArena Coding1385—
BigCodeBench Complete—48.7%

Agentic & Tool Use Not comparable

Nova 2 Lite: 24.1 (#121), phi-3-medium 14B: —

Agentic & Tool Use benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
Berkeley Function Calling Leaderboard27.1%—

Reasoning Not comparable

Nova 2 Lite: 27.5 (#118), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
LMArena Hard Prompts1364—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
Epoch Capabilities Index—121.23
HellaSwag—82.4%
WinoGrande—81.5%

Math Nova 2 Lite leads

Nova 2 Lite: 37.5 (#156), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
LMArena Math1359—
MATH Level 5—17.6%

Knowledge Nova 2 Lite leads

Nova 2 Lite: 43.0 (#94), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
GPQA Diamond—27.6%
Vectara Hallucination Rate5.1%—
LMArena Expert1358—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

Nova 2 Lite: 47.1 (#153), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
LMArena Non-English1337—
LMArena Chinese1364—
LMArena French1381—
LMArena German1343—
LMArena Japanese1271—
LMArena Korean1284—
LMArena Russian1343—
LMArena Spanish1373—

Instruction Following Not comparable

Nova 2 Lite: 70.5 (#161), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
LMArena Instruction Following1335—

Long Context Not comparable

Nova 2 Lite: 40.6 (#150), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
LMArena Longer Query1335—

Writing & Preference Not comparable

Nova 2 Lite: 53.9 (#154), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkNova 2 Litephi-3-medium 14B
LMArena Text1362—
LMArena Creative Writing1291—
LMArena Multi-Turn1338—

Frequently asked questions

Is Nova 2 Lite better than phi-3-medium 14B?

Nova 2 Lite is the stronger model overall, scoring 39.7 to 29.7 on the Noometry Index.

Is Nova 2 Lite or phi-3-medium 14B better for coding?

Nova 2 Lite scores higher on coding benchmarks: 40.7 versus 36.8 in the Noometry coding category.

How many benchmarks do Nova 2 Lite and phi-3-medium 14B share?

0 benchmarks have published results for both models. Nova 2 Lite has 19 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper