Model comparison

Llama 3.3 Nemotron 49b Super v1 vs Llama2 70b Steerlm Chat

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 31.8 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 3.3 Nemotron 49b Super v1 scores higher in 6 categories and Llama2 70b Steerlm Chat in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.3 Nemotron 49b Super v1 leads 50.8 to 31.6.

Side by side

Llama 3.3 Nemotron 49b Super v1 and Llama2 70b Steerlm Chat specifications
Llama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
ProviderNVIDIANVIDIA
Noometry Index40.131.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked109

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 37.9 (#186), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Coding12961025

Reasoning Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 26.2 (#135), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Hard Prompts13111047

Math Not comparable

Llama 3.3 Nemotron 49b Super v1: —, Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Math—1072

Multilingual Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 41.1 (#211), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Non-English12531063
LMArena Chinese1277—
LMArena Russian1269—

Instruction Following Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 68.3 (#189), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Instruction Following12931060

Long Context Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 39.5 (#176), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Longer Query1299998

Writing & Preference Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 50.8 (#179), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Llama2 70b Steerlm Chat
LMArena Text13081098
LMArena Creative Writing12881091
LMArena Multi-Turn13151058

Frequently asked questions

Is Llama 3.3 Nemotron 49b Super v1 better than Llama2 70b Steerlm Chat?

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 31.8 on the Noometry Index.

Is Llama 3.3 Nemotron 49b Super v1 or Llama2 70b Steerlm Chat better for coding?

Llama 3.3 Nemotron 49b Super v1 scores higher on coding benchmarks: 37.9 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama 3.3 Nemotron 49b Super v1 and Llama2 70b Steerlm Chat share?

8 benchmarks have published results for both models. Llama 3.3 Nemotron 49b Super v1 has 10 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper