Model comparison

Nvidia Llama 3.3 Nemotron Super 49b v1.5 vs Phi 3 Small 8k Instruct

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 29.3 on the Noometry Index.

Last verified . 12 shared benchmarks.

Summary

  • They share 12 benchmarks with published results for both. Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher in 8 categories and Phi 3 Small 8k Instruct in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 53.1 to 31.1.

Side by side

Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Phi 3 Small 8k Instruct specifications
Nvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
ProviderNVIDIAMicrosoft
Noometry Index40.329.3
Released2025-07-252024-04-23
WeightsOpenOpen
Context window131K—
Max output131K—
Input $ / M tokens$0.40—
Output $ / M tokens$0.40—
Results tracked1232

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Coding13551101
LiveBench Coding—20.3%

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Hard Prompts13361100
LiveBench Reasoning—15.9%
LiveBench Data Analysis—30.3%
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
HellaSwag—77%
LiveBench—24%
WinoGrande—81.5%

Math Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Math13921151
LiveBench Math—17.6%

Knowledge Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Expert13301067
ARC (AI2) Challenge—90.7%
MMLU—75.7%
OpenBookQA—88%
TriviaQA—58.1%

Multilingual Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Non-English13161058
LMArena Japanese1300966
LMArena Russian13321111
LMArena Chinese—1061
LMArena French—1135
LMArena German—1080
LMArena Korean—894
LMArena Spanish—1111

Instruction Following Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Instruction Following12991087
LiveBench Instruction Following—47.2%

Long Context Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Longer Query13151088

Writing & Preference Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Phi 3 Small 8k Instruct
LMArena Text13381110
LMArena Creative Writing13071083
LMArena Multi-Turn13341068
LiveBench Language—12.9%

Frequently asked questions

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 better than Phi 3 Small 8k Instruct?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 29.3 on the Noometry Index.

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 or Phi 3 Small 8k Instruct better for coding?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher on coding benchmarks: 39.8 versus 27.9 in the Noometry coding category.

How many benchmarks do Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Phi 3 Small 8k Instruct share?

12 benchmarks have published results for both models. Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper