Model comparison

Llama 3.1 Nemotron 70b Instruct vs Llama 3.3 Nemotron 49b Super v1

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 37.6 on the Noometry Index.

Last verified . 10 shared benchmarks.

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron 70b Instruct scores higher in 0 categories and Llama 3.3 Nemotron 49b Super v1 in 6 categories; 5 gaps are clear of the uncertainty.

Side by side

Llama 3.1 Nemotron 70b Instruct and Llama 3.3 Nemotron 49b Super v1 specifications
Llama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
ProviderNVIDIANVIDIA
Noometry Index37.640.1
Released2024-12-18—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.1 Nemotron 70b Instruct: 35.9 (#216), Llama 3.3 Nemotron 49b Super v1: 37.9 (#186)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Coding12721296
BigCodeBench Instruct38.7%—
BigCodeBench Complete48.2%—

Reasoning Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.1 Nemotron 70b Instruct: 25.0 (#152), Llama 3.3 Nemotron 49b Super v1: 26.2 (#135)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Hard Prompts12661311

Math Not comparable

Llama 3.1 Nemotron 70b Instruct: 35.5 (#182), Llama 3.3 Nemotron 49b Super v1: —

Math benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Math1271—

Knowledge Not comparable

Llama 3.1 Nemotron 70b Instruct: 34.1 (#199), Llama 3.3 Nemotron 49b Super v1: —

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Expert1242—

Multilingual Too close to call

Llama 3.1 Nemotron 70b Instruct: 40.5 (#217), Llama 3.3 Nemotron 49b Super v1: 41.1 (#211)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Non-English12451253
LMArena Chinese12631277
LMArena Russian12271269

Instruction Following Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.1 Nemotron 70b Instruct: 65.9 (#213), Llama 3.3 Nemotron 49b Super v1: 68.3 (#189)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Instruction Following12521293

Long Context Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.1 Nemotron 70b Instruct: 37.6 (#215), Llama 3.3 Nemotron 49b Super v1: 39.5 (#176)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Longer Query12381299

Writing & Preference Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.1 Nemotron 70b Instruct: 48.4 (#203), Llama 3.3 Nemotron 49b Super v1: 50.8 (#179)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructLlama 3.3 Nemotron 49b Super v1
LMArena Text12831308
LMArena Creative Writing12691288
LMArena Multi-Turn12751315

Frequently asked questions

Is Llama 3.1 Nemotron 70b Instruct better than Llama 3.3 Nemotron 49b Super v1?

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 37.6 on the Noometry Index.

Is Llama 3.1 Nemotron 70b Instruct or Llama 3.3 Nemotron 49b Super v1 better for coding?

Llama 3.3 Nemotron 49b Super v1 scores higher on coding benchmarks: 37.9 versus 35.9 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron 70b Instruct and Llama 3.3 Nemotron 49b Super v1 share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron 70b Instruct has 14 scored results on Noometry and Llama 3.3 Nemotron 49b Super v1 has 10.

Related comparisons

Go deeper