Model comparison

DeepSeek-R1-Distill-Qwen-1.5B vs Llama 3.1 Nemotron Ultra 253b v1

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 26.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in coding, where Llama 3.1 Nemotron Ultra 253b v1 leads 38.4 to 21.8.

Side by side

DeepSeek-R1-Distill-Qwen-1.5B and Llama 3.1 Nemotron Ultra 253b v1 specifications
DeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
ProviderDeepSeekNVIDIA
Noometry Index26.136.7
Released2025-01-20—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked511

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Qwen-1.5B: 21.8 (#336), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
BigCodeBench Instruct7%—
LMArena Coding—1312
BigCodeBench Complete7.9%—

Agentic & Tool Use Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard—10%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Qwen-1.5B: 19.2 (#262), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
Chess Puzzles0%—
LMArena Hard Prompts—1316

Math Llama 3.1 Nemotron Ultra 253b v1 leads

DeepSeek-R1-Distill-Qwen-1.5B: 23.0 (#274), Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
OTIS Mock AIME 2024-202521.4%—
LMArena Math—1360

Knowledge Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: 16.0 (#290), Llama 3.1 Nemotron Ultra 253b v1: —

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
GPQA Diamond33.6%—

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
LMArena Non-English—1282
LMArena Russian—1284

Instruction Following Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
LMArena Instruction Following—1308

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query—1299

Writing & Preference Not comparable

DeepSeek-R1-Distill-Qwen-1.5B: —, Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-1.5BLlama 3.1 Nemotron Ultra 253b v1
LMArena Text—1320
LMArena Creative Writing—1314
LMArena Multi-Turn—1317

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-1.5B better than Llama 3.1 Nemotron Ultra 253b v1?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 26.1 on the Noometry Index.

Is DeepSeek-R1-Distill-Qwen-1.5B or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 21.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Qwen-1.5B and Llama 3.1 Nemotron Ultra 253b v1 share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-1.5B has 5 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper