Model comparison

Dolly 2.0-12b vs Llama 3.3 Nemotron 49b Super v1

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 25.5 on the Noometry Index.

Last verified . 8 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Dolly 2.0-12b scores higher in 0 categories and Llama 3.3 Nemotron 49b Super v1 in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.3 Nemotron 49b Super v1 leads 50.8 to 15.2.

Side by side

Dolly 2.0-12b and Llama 3.3 Nemotron 49b Super v1 specifications
Dolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
ProviderDatabricksNVIDIA
Noometry Index25.540.1
Released2023-04-11—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.3 Nemotron 49b Super v1 leads

Dolly 2.0-12b: 23.4 (#332), Llama 3.3 Nemotron 49b Super v1: 37.9 (#186)

Coding benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Coding7761296

Reasoning Llama 3.3 Nemotron 49b Super v1 leads

Dolly 2.0-12b: 15.3 (#316), Llama 3.3 Nemotron 49b Super v1: 26.2 (#135)

Reasoning benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Hard Prompts8041311
Epoch Capabilities Index89.67—
HellaSwag70.8%—
PIQA75.4%—
WinoGrande61.8%—

Math Not comparable

Dolly 2.0-12b: 27.3 (#251), Llama 3.3 Nemotron 49b Super v1: —

Math benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Math871—

Knowledge Not comparable

Dolly 2.0-12b: —, Llama 3.3 Nemotron 49b Super v1: —

Knowledge benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
ARC (AI2) Challenge39.6%—
BoolQ56.3%—
MMLU26.2%—
OpenBookQA39.2%—

Multilingual Llama 3.3 Nemotron 49b Super v1 leads

Dolly 2.0-12b: 17.4 (#296), Llama 3.3 Nemotron 49b Super v1: 41.1 (#211)

Multilingual benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Non-English8361253
LMArena Chinese8361277
LMArena Russian—1269

Instruction Following Llama 3.3 Nemotron 49b Super v1 leads

Dolly 2.0-12b: 38.7 (#304), Llama 3.3 Nemotron 49b Super v1: 68.3 (#189)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Instruction Following8141293

Long Context Not comparable

Dolly 2.0-12b: —, Llama 3.3 Nemotron 49b Super v1: 39.5 (#176)

Long Context benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Longer Query—1299

Writing & Preference Llama 3.3 Nemotron 49b Super v1 leads

Dolly 2.0-12b: 15.2 (#311), Llama 3.3 Nemotron 49b Super v1: 50.8 (#179)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bLlama 3.3 Nemotron 49b Super v1
LMArena Text8511308
LMArena Creative Writing8641288
LMArena Multi-Turn7401315

Frequently asked questions

Is Dolly 2.0-12b better than Llama 3.3 Nemotron 49b Super v1?

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 25.5 on the Noometry Index.

Is Dolly 2.0-12b or Llama 3.3 Nemotron 49b Super v1 better for coding?

Llama 3.3 Nemotron 49b Super v1 scores higher on coding benchmarks: 37.9 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Llama 3.3 Nemotron 49b Super v1 share?

8 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Llama 3.3 Nemotron 49b Super v1 has 10.

Related comparisons

Go deeper