Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Llama 4 Scout

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 27.7 on the Noometry Index.

Last verified . 11 shared benchmarks.

Llama 4 Scout Meta

27.7

Rank #330 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 7 categories and Llama 4 Scout in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Llama 3.1 Nemotron Ultra 253b v1 leads 38.4 to 20.2.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 10% for Llama 3.1 Nemotron Ultra 253b v1 and 28.1% for Llama 4 Scout.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Llama 4 Scout specifications
Llama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
ProviderNVIDIAMeta
Noometry Index36.727.7
Released—2025-04-05
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.30
Results tracked1143

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Llama 4 Scout: 20.2 (#339)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Coding13121286
SWE-bench Verified (bash only)—9.1%
SciCode—17%
BigCodeBench Complete—43.1%

Agentic & Tool Use Llama 4 Scout leads

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Llama 4 Scout: 24.6 (#119)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
Berkeley Function Calling Leaderboard10%28.1%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Llama 4 Scout: 9.1 (#345)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Hard Prompts13161266
ARC-AGI-2—0%
Kagi LLM Benchmark—36.9%
ARC-AGI-1—0.5%
CritPt—0%
DTBench—57.9%
LMCA—12%
Epoch Capabilities Index—129.64
ForecastBench—57.5

Math Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Llama 4 Scout: 19.6 (#286)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Math13601287
OTIS Mock AIME 2024-2025—7.8%
Omni-MATH—37.3%
MATH Level 5—62.3%
FrontierMath (Feb 2025 set)—0%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Llama 4 Scout: 31.9 (#217)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
GPQA Diamond—51.8%
MMLU-Pro—74.2%
Vectara Hallucination Rate—7.7%
GPQA (HELM)—50.7%
LMArena Expert—1235

Multimodal Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Llama 4 Scout: 32.2 (#102)

Multimodal benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Vision—1118
SpatialViz-Bench—34.2%

Multilingual Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Llama 4 Scout: 41.0 (#212)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Non-English12821252
LMArena Russian12841263
LMArena Chinese—1255
LMArena French—1282
LMArena German—1272
LMArena Japanese—1206
LMArena Korean—1207
LMArena Spanish—1278

Instruction Following Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Llama 4 Scout: 65.8 (#217)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Instruction Following13081248
IFEval—81.8%

Long Context Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Llama 4 Scout: 27.5 (#294)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Longer Query12991265
Fiction.LiveBench—36%

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Llama 4 Scout: 37.0 (#261)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Llama 4 Scout
LMArena Text13201279
LMArena Creative Writing13141249
LMArena Multi-Turn13171280
EQ-Bench Creative Writing—783
WildBench—78%

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Llama 4 Scout?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 27.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Llama 4 Scout better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 20.2 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Llama 4 Scout share?

11 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Llama 4 Scout has 43.

Related comparisons

Go deeper