Model comparison

Dolly 2.0-12b vs Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 25.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Nemotron 3.5 Lightning NVIDIA

40.0

Rank #155 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Dolly 2.0-12b scores higher in 0 categories and Nemotron 3.5 Lightning in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Nemotron 3.5 Lightning leads 48.5 to 15.2.

Side by side

Dolly 2.0-12b and Nemotron 3.5 Lightning specifications
Dolly 2.0-12bNemotron 3.5 Lightning
ProviderDatabricksNVIDIA
Noometry Index25.540.0
Released2023-04-112026-08-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.20
Results tracked1718

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3.5 Lightning leads

Dolly 2.0-12b: 23.4 (#332), Nemotron 3.5 Lightning: 40.4 (#141)

Coding benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Coding7761375

Reasoning Nemotron 3.5 Lightning leads

Dolly 2.0-12b: 15.3 (#316), Nemotron 3.5 Lightning: 26.8 (#127)

Reasoning benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Hard Prompts8041337
Epoch Capabilities Index89.67—
HellaSwag70.8%—
PIQA75.4%—
WinoGrande61.8%—

Math Nemotron 3.5 Lightning leads

Dolly 2.0-12b: 27.3 (#251), Nemotron 3.5 Lightning: 37.5 (#155)

Math benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Math8711359

Knowledge Not comparable

Dolly 2.0-12b: —, Nemotron 3.5 Lightning: 37.5 (#154)

Knowledge benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Expert—1356
ARC (AI2) Challenge39.6%—
BoolQ56.3%—
MMLU26.2%—
OpenBookQA39.2%—

Multilingual Nemotron 3.5 Lightning leads

Dolly 2.0-12b: 17.4 (#296), Nemotron 3.5 Lightning: 44.0 (#180)

Multilingual benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Non-English8361295
LMArena Chinese8361359
LMArena French—1366
LMArena German—1282
LMArena Japanese—1206
LMArena Korean—1238
LMArena Russian—1253
LMArena Spanish—1345

Instruction Following Nemotron 3.5 Lightning leads

Dolly 2.0-12b: 38.7 (#304), Nemotron 3.5 Lightning: 69.6 (#170)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Instruction Following8141318

Long Context Not comparable

Dolly 2.0-12b: —, Nemotron 3.5 Lightning: 39.9 (#165)

Long Context benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Longer Query—1314

Writing & Preference Nemotron 3.5 Lightning leads

Dolly 2.0-12b: 15.2 (#311), Nemotron 3.5 Lightning: 48.5 (#201)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bNemotron 3.5 Lightning
LMArena Text8511327
LMArena Creative Writing8641254
LMArena Multi-Turn7401328
EQ-Bench Creative Writing—1280

Frequently asked questions

Is Dolly 2.0-12b better than Nemotron 3.5 Lightning?

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 25.5 on the Noometry Index.

Is Dolly 2.0-12b or Nemotron 3.5 Lightning better for coding?

Nemotron 3.5 Lightning scores higher on coding benchmarks: 40.4 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Nemotron 3.5 Lightning share?

9 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Nemotron 3.5 Lightning has 18.

Related comparisons

Go deeper