Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 36.7 on the Noometry Index.

Last verified . 9 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 2 categories and Mercury 2 in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Llama 3.1 Nemotron Ultra 253b v1 leads 38.4 to 33.5.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Mercury 2 specifications
Llama 3.1 Nemotron Ultra 253b v1Mercury 2
ProviderNVIDIAInception
Noometry Index36.739.1
Released—2026-02-20
WeightsOpenProprietary
Context window—128K
Max output—50K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked1117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Coding13121391
LMArena WebDev—1171
SciCode—38.7%
WeirdML—43.2%
ALE-Bench—785.58

Agentic & Tool Use Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Mercury 2: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
Berkeley Function Calling Leaderboard10%—

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Hard Prompts13161362
CritPt—0.8%

Math Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Mercury 2: —

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Math1360—

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
Vectara Hallucination Rate—12.3%
LMArena Expert—1358

Multilingual Mercury 2 leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Non-English12821331
LMArena Russian12841304
LMArena Chinese—1417

Instruction Following Mercury 2 leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Instruction Following13081329

Long Context Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Longer Query12991330

Writing & Preference Mercury 2 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mercury 2
LMArena Text13201355
LMArena Creative Writing13141289
LMArena Multi-Turn13171358

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 36.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Mercury 2 better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 33.5 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Mercury 2 share?

9 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper