Model comparison

DeepSeek-R1-Distill-Llama-70B vs Llama 3.1 Nemotron Ultra 253b v1

DeepSeek-R1-Distill-Llama-70B is the stronger model overall, scoring 37.8 to 36.7 on the Noometry Index.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in writing & preference, where Llama 3.1 Nemotron Ultra 253b v1 leads 52.2 to 49.0.

Side by side

DeepSeek-R1-Distill-Llama-70B and Llama 3.1 Nemotron Ultra 253b v1 specifications
DeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
ProviderDeepSeekNVIDIA
Noometry Index37.836.7
Released2025-01-20—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1311

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Llama-70B: 36.8 (#202), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
BigCodeBench Instruct35.3%—
LiveBench Coding51.6%—
LMArena Coding—1312
BigCodeBench Complete49.9%—

Agentic & Tool Use Not comparable

DeepSeek-R1-Distill-Llama-70B: —, Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard—10%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Llama-70B: 24.9 (#156), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
Kagi LLM Benchmark52.3%—
LiveBench Reasoning67.6%—
LMArena Hard Prompts—1316
LiveBench Data Analysis55.9%—
LiveBench54.5%—

Math Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Llama-70B: 36.0 (#176), Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
OTIS Mock AIME 2024-202551.4%—
LiveBench Math58.1%—
LMArena Math—1360
MATH Level 589.9%—

Knowledge Not comparable

DeepSeek-R1-Distill-Llama-70B: 30.7 (#225), Llama 3.1 Nemotron Ultra 253b v1: —

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
GPQA Diamond55.7%—

Multilingual Not comparable

DeepSeek-R1-Distill-Llama-70B: —, Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
LMArena Non-English—1282
LMArena Russian—1284

Instruction Following Too close to call

DeepSeek-R1-Distill-Llama-70B: 68.2 (#190), Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
LiveBench Instruction Following69.9%—
LMArena Instruction Following—1308

Long Context Not comparable

DeepSeek-R1-Distill-Llama-70B: —, Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query—1299

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Llama-70B: 49.0 (#194), Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BLlama 3.1 Nemotron Ultra 253b v1
LMArena Text—1320
LMArena Creative Writing—1314
LMArena Multi-Turn—1317
LiveBench Language23.8%—

Frequently asked questions

Is DeepSeek-R1-Distill-Llama-70B better than Llama 3.1 Nemotron Ultra 253b v1?

DeepSeek-R1-Distill-Llama-70B is the stronger model overall, scoring 37.8 to 36.7 on the Noometry Index.

Is DeepSeek-R1-Distill-Llama-70B or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 36.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Llama-70B and Llama 3.1 Nemotron Ultra 253b v1 share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Llama-70B has 13 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper