Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Llama 3.2 1B

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 20.1 on the Noometry Index.

Last verified . 11 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 8 categories and Llama 3.2 1B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.1 Nemotron Ultra 253b v1 leads 52.2 to 21.3.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Llama 3.2 1B specifications
Llama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
ProviderNVIDIAMeta
Noometry Index36.720.1
Released—2024-09-24
WeightsOpenOpen
Context window—60K
Max output—54K
Input $ / M tokens—$0.027
Output $ / M tokens—$0.20
Results tracked1122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Coding13121070
BigCodeBench Instruct—8.2%
BigCodeBench Complete—11.3%

Agentic & Tool Use Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
Berkeley Function Calling Leaderboard10%10.8%
BALROG—6.6%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Hard Prompts13161044
Chess Puzzles—0%
Epoch Capabilities Index—101.99

Math Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Math13601086
OTIS Mock AIME 2024-2025—0.6%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
GPQA Diamond—23.9%
LMArena Expert—1007

Multilingual Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Non-English1282973
LMArena Russian1284941
LMArena Chinese—959
LMArena German—1014

Instruction Following Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Instruction Following13081031

Long Context Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Longer Query12991050

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 3.2 1B
LMArena Text13201055
LMArena Creative Writing13141033
LMArena Multi-Turn13171030
EQ-Bench Creative Writing—200

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Llama 3.2 1B?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 20.1 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Llama 3.2 1B better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 21.1 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Llama 3.2 1B share?

11 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper