Model comparison

DeepSeek-V2.5 (Sep 2024) vs Nemotron 3 Ultra

Nemotron 3 Ultra is the stronger model overall, scoring 42.5 to 37.6 on the Noometry Index.

Last verified . 16 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Nemotron 3 Ultra NVIDIA

42.5

Rank #113 Confirmed

Summary

  • They share 16 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 1 category and Nemotron 3 Ultra in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Nemotron 3 Ultra leads 52.5 to 34.8.

Side by side

DeepSeek-V2.5 (Sep 2024) and Nemotron 3 Ultra specifications
DeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
ProviderDeepSeekNVIDIA
Noometry Index37.642.5
Released2024-09-062026-06-04
WeightsOpenOpen
Context window—262K
Max output—128K
Input $ / M tokens—$0.50
Output $ / M tokens—$2.20
Results tracked2230

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Nemotron 3 Ultra: 38.1 (#182)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Coding13091468
FrontierCode—13.6%
Aider Polyglot17.8%—
SciCode—40.3%
WeirdML—43.5%
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Agentic & Tool Use Not comparable

DeepSeek-V2.5 (Sep 2024): —, Nemotron 3 Ultra: 23.4 (#126)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
APEX-Agents—22.7%

Reasoning Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Nemotron 3 Ultra: 31.0 (#82)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Hard Prompts12891452
CritPt—3.1%
Chess Puzzles—12%
Mystery Game Puzzles—20%
DTBench—90.1%
LMCA—36.9%
Epoch Capabilities Index—146.17

Math Too close to call

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Nemotron 3 Ultra: 35.3 (#187)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Math12881457
OTIS Mock AIME 2024-2025—86.7%
ProofBench—2%

Knowledge Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Nemotron 3 Ultra: 52.5 (#61)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Expert12661472
GPQA Diamond—85.4%

Multilingual Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Nemotron 3 Ultra: 53.5 (#64)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Non-English12731427
LMArena Chinese13181497
LMArena French12891472
LMArena German12581471
LMArena Korean12091386
LMArena Russian12891417
LMArena Spanish12481454
LMArena Japanese1228—

Instruction Following Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Nemotron 3 Ultra: 74.7 (#88)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Instruction Following12801418

Long Context Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Nemotron 3 Ultra: 43.9 (#81)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Longer Query13011437

Writing & Preference Nemotron 3 Ultra leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Nemotron 3 Ultra: 66.3 (#36)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Nemotron 3 Ultra
LMArena Text12941445
LMArena Creative Writing12851398
LMArena Multi-Turn12971411
EQ-Bench Creative Writing—1692

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Nemotron 3 Ultra?

Nemotron 3 Ultra is the stronger model overall, scoring 42.5 to 37.6 on the Noometry Index.

Is DeepSeek-V2.5 (Sep 2024) or Nemotron 3 Ultra better for coding?

Nemotron 3 Ultra scores higher on coding benchmarks: 38.1 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Nemotron 3 Ultra share?

16 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Nemotron 3 Ultra has 30.

Related comparisons

Go deeper