Model comparison

Llama 3.2 3B vs Nemotron 3 Super

Nemotron 3 Super is the stronger model overall, scoring 40.1 to 28.9 on the Noometry Index.

Last verified . 13 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Nemotron 3 Super NVIDIA

40.1

Rank #153 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.2 3B scores higher in 1 category and Nemotron 3 Super in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Nemotron 3 Super leads 55.9 to 24.7.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.08 / $0.45 for Nemotron 3 Super.
  • Nemotron 3 Super accepts more context: 262K tokens versus 131K.

Side by side

Llama 3.2 3B and Nemotron 3 Super specifications
Llama 3.2 3BNemotron 3 Super
ProviderMetaNVIDIA
Noometry Index28.940.1
Released2024-09-242026-03-11
WeightsOpenOpen
Context window131K262K
Max output118K131K
Input $ / M tokens$0.05$0.08
Output $ / M tokens$0.33$0.45
Results tracked1821

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3 Super leads

Llama 3.2 3B: 27.6 (#319), Nemotron 3 Super: 39.4 (#158)

Coding benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Coding10981403
SciCode—36%
WeirdML—38%
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—
ALE-Bench—213.9

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Nemotron 3 Super: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Nemotron 3 Super: 18.7 (#276)

Reasoning benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Hard Prompts10951386
NYT Connections (extended)—15.4%
CritPt—3.1%

Math Nemotron 3 Super leads

Llama 3.2 3B: 32.4 (#214), Nemotron 3 Super: 39.6 (#101)

Math benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Math11261380
MathArena Final-Answer Competitions—60.4%

Knowledge Nemotron 3 Super leads

Llama 3.2 3B: 29.7 (#235), Nemotron 3 Super: 38.9 (#139)

Knowledge benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Expert10901398

Multilingual Nemotron 3 Super leads

Llama 3.2 3B: 26.2 (#281), Nemotron 3 Super: 48.4 (#144)

Multilingual benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Non-English10191355
LMArena Chinese10171432
LMArena German10561341
LMArena Russian9491336
LMArena French—1405
LMArena Spanish—1417

Instruction Following Nemotron 3 Super leads

Llama 3.2 3B: 56.0 (#275), Nemotron 3 Super: 71.2 (#154)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Instruction Following10891347

Long Context Nemotron 3 Super leads

Llama 3.2 3B: 33.4 (#261), Nemotron 3 Super: 41.5 (#139)

Long Context benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Longer Query11001362

Writing & Preference Nemotron 3 Super leads

Llama 3.2 3B: 24.7 (#307), Nemotron 3 Super: 55.9 (#140)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BNemotron 3 Super
LMArena Text11101378
LMArena Creative Writing10941317
LMArena Multi-Turn11051369
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Nemotron 3 Super?

Nemotron 3 Super is the stronger model overall, scoring 40.1 to 28.9 on the Noometry Index.

Which is cheaper, Llama 3.2 3B or Nemotron 3 Super?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Nemotron 3 Super lists at $0.08 and $0.45.

Is Llama 3.2 3B or Nemotron 3 Super better for coding?

Nemotron 3 Super scores higher on coding benchmarks: 39.4 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Nemotron 3 Super does, with 262K tokens against 131K.

How many benchmarks do Llama 3.2 3B and Nemotron 3 Super share?

13 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Nemotron 3 Super has 21.

Related comparisons

Go deeper