Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Llama 3.3 Nemotron 49b Super v1

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 36.7 on the Noometry Index.

Last verified . 9 shared benchmarks.

Summary

  • They share 9 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 5 categories and Llama 3.3 Nemotron 49b Super v1 in 1 category; 2 gaps are clear of the uncertainty.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Llama 3.3 Nemotron 49b Super v1 specifications
Llama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
ProviderNVIDIANVIDIA
Noometry Index36.740.1
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1110

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Llama 3.3 Nemotron 49b Super v1: 37.9 (#186)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Coding13121296

Agentic & Tool Use Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Llama 3.3 Nemotron 49b Super v1: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
Berkeley Function Calling Leaderboard10%—

Reasoning Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Llama 3.3 Nemotron 49b Super v1: 26.2 (#135)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Hard Prompts13161311

Math Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Llama 3.3 Nemotron 49b Super v1: —

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Math1360—

Multilingual Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Llama 3.3 Nemotron 49b Super v1: 41.1 (#211)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Non-English12821253
LMArena Russian12841269
LMArena Chinese—1277

Instruction Following Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Llama 3.3 Nemotron 49b Super v1: 68.3 (#189)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Instruction Following13081293

Long Context Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Llama 3.3 Nemotron 49b Super v1: 39.5 (#176)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Longer Query12991299

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Llama 3.3 Nemotron 49b Super v1: 50.8 (#179)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.3 Nemotron 49b Super v1
LMArena Text13201308
LMArena Creative Writing13141288
LMArena Multi-Turn13171315

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Llama 3.3 Nemotron 49b Super v1?

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 36.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Llama 3.3 Nemotron 49b Super v1 better for coding?

They score almost the same on coding (38.4 vs 37.9); test both on your own repository before choosing.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Llama 3.3 Nemotron 49b Super v1 share?

9 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Llama 3.3 Nemotron 49b Super v1 has 10.

Related comparisons

Go deeper