Model comparison

Llama 3.1 Nemotron 51b Instruct vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 35.9 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 3.1 Nemotron 51b Instruct scores higher in 1 category and Mercury in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.1 Nemotron 51b Instruct leads 23.5 to 17.5.
  • Llama 3.1 Nemotron 51b Instruct has downloadable open weights; the other is API-only.

Side by side

Llama 3.1 Nemotron 51b Instruct and Mercury specifications
Llama 3.1 Nemotron 51b InstructMercury
ProviderNVIDIAInception
Noometry Index35.937.6
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked129

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Llama 3.1 Nemotron 51b Instruct: 35.6 (#222), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Coding12231322

Reasoning Llama 3.1 Nemotron 51b Instruct leads

Llama 3.1 Nemotron 51b Instruct: 23.5 (#177), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Hard Prompts12031285
Kagi LLM Benchmark—21.6%

Math Not comparable

Llama 3.1 Nemotron 51b Instruct: 34.6 (#193), Mercury: —

Math benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Math1230—

Knowledge Not comparable

Llama 3.1 Nemotron 51b Instruct: 31.9 (#218), Mercury: —

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Expert1167—

Multilingual Mercury leads

Llama 3.1 Nemotron 51b Instruct: 36.1 (#241), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Non-English11811260
LMArena Chinese1180—
LMArena Russian1187—

Instruction Following Mercury leads

Llama 3.1 Nemotron 51b Instruct: 62.9 (#233), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Instruction Following12011239

Long Context Mercury leads

Llama 3.1 Nemotron 51b Instruct: 36.5 (#230), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Longer Query12051266

Writing & Preference Mercury leads

Llama 3.1 Nemotron 51b Instruct: 43.4 (#229), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMercury
LMArena Text12281282
LMArena Creative Writing12131191
LMArena Multi-Turn12271282

Frequently asked questions

Is Llama 3.1 Nemotron 51b Instruct better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 35.9 on the Noometry Index.

Is Llama 3.1 Nemotron 51b Instruct or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 35.6 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron 51b Instruct and Mercury share?

8 benchmarks have published results for both models. Llama 3.1 Nemotron 51b Instruct has 12 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper