Model comparison

Nemotron 3 Ultra vs phi-3-medium 14B

Nemotron 3 Ultra is the stronger model overall, scoring 42.5 to 29.7 on the Noometry Index.

Last verified . 2 shared benchmarks.

Nemotron 3 Ultra NVIDIA

42.5

Rank #113 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. Nemotron 3 Ultra scores higher in 3 categories and phi-3-medium 14B in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Nemotron 3 Ultra leads 52.5 to 9.1.
  • The biggest single-benchmark swing is GPQA Diamond: 85.4% for Nemotron 3 Ultra and 27.6% for phi-3-medium 14B.

Side by side

Nemotron 3 Ultra and phi-3-medium 14B specifications
Nemotron 3 Ultraphi-3-medium 14B
ProviderNVIDIAMicrosoft
Noometry Index42.529.7
Released2026-06-042024-04-23
WeightsOpenOpen
Context window262K—
Max output128K—
Input $ / M tokens$0.60—
Output $ / M tokens$2.40—
Results tracked3013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3 Ultra leads

Nemotron 3 Ultra: 38.1 (#182), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
FrontierCode13.6%—
SciCode40.3%—
WeirdML43.5%—
BigCodeBench Instruct—37.6%
LMArena Coding1468—
BigCodeBench Complete—48.7%

Agentic & Tool Use Not comparable

Nemotron 3 Ultra: 23.4 (#126), phi-3-medium 14B: —

Agentic & Tool Use benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
APEX-Agents22.7%—

Reasoning Not comparable

Nemotron 3 Ultra: 31.0 (#82), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
Epoch Capabilities Index146.17121.23
CritPt3.1%—
Chess Puzzles12%—
LMArena Hard Prompts1452—
Mystery Game Puzzles20%—
DTBench90.1%—
LMCA36.9%—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
HellaSwag—82.4%
WinoGrande—81.5%

Math Nemotron 3 Ultra leads

Nemotron 3 Ultra: 35.3 (#187), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
OTIS Mock AIME 2024-202586.7%—
ProofBench2%—
LMArena Math1457—
MATH Level 5—17.6%

Knowledge Nemotron 3 Ultra leads

Nemotron 3 Ultra: 52.5 (#61), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
GPQA Diamond85.4%27.6%
LMArena Expert1472—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

Nemotron 3 Ultra: 53.5 (#64), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
LMArena Non-English1427—
LMArena Chinese1497—
LMArena French1472—
LMArena German1471—
LMArena Korean1386—
LMArena Russian1417—
LMArena Spanish1454—

Instruction Following Not comparable

Nemotron 3 Ultra: 74.7 (#88), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
LMArena Instruction Following1418—

Long Context Not comparable

Nemotron 3 Ultra: 43.9 (#81), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
LMArena Longer Query1437—

Writing & Preference Not comparable

Nemotron 3 Ultra: 66.3 (#36), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkNemotron 3 Ultraphi-3-medium 14B
LMArena Text1445—
LMArena Creative Writing1398—
EQ-Bench Creative Writing1692—
LMArena Multi-Turn1411—

Frequently asked questions

Is Nemotron 3 Ultra better than phi-3-medium 14B?

Nemotron 3 Ultra is the stronger model overall, scoring 42.5 to 29.7 on the Noometry Index.

Is Nemotron 3 Ultra or phi-3-medium 14B better for coding?

Nemotron 3 Ultra scores higher on coding benchmarks: 38.1 versus 36.8 in the Noometry coding category.

How many benchmarks do Nemotron 3 Ultra and phi-3-medium 14B share?

2 benchmarks have published results for both models. Nemotron 3 Ultra has 30 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper