Model comparison

Dolly 2.0-12b vs Llama 13b

Dolly 2.0-12b is the stronger model overall, scoring 25.5 to 24.4 on the Noometry Index.

Last verified . 16 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Dolly 2.0-12b scores higher in 6 categories and Llama 13b in 0 categories; 4 gaps are clear of the uncertainty.

Side by side

Dolly 2.0-12b and Llama 13b specifications
Dolly 2.0-12bLlama 13b
ProviderDatabricksMeta
Noometry Index25.524.4
Released2023-04-112023-02-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1721

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Dolly 2.0-12b leads

Dolly 2.0-12b: 23.4 (#332), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkDolly 2.0-12bLlama 13b
LMArena Coding776683

Reasoning Dolly 2.0-12b leads

Dolly 2.0-12b: 15.3 (#316), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkDolly 2.0-12bLlama 13b
LMArena Hard Prompts804728
Epoch Capabilities Index89.67100.58
HellaSwag70.8%79.2%
PIQA75.4%80.1%
WinoGrande61.8%73%
BIG-Bench Hard—37.9%
LAMBADA—75.2%

Math Too close to call

Dolly 2.0-12b: 27.3 (#251), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkDolly 2.0-12bLlama 13b
LMArena Math871838
GSM8K—20.6%

Knowledge Not comparable

Dolly 2.0-12b: —, Llama 13b: —

Knowledge benchmarks
BenchmarkDolly 2.0-12bLlama 13b
ARC (AI2) Challenge39.6%52.7%
BoolQ56.3%78.7%
MMLU26.2%47.7%
OpenBookQA39.2%56.4%
TriviaQA—77.9%

Multimodal Not comparable

Dolly 2.0-12b: —, Llama 13b: —

Multimodal benchmarks
BenchmarkDolly 2.0-12bLlama 13b
ScienceQA—43.3%

Multilingual Too close to call

Dolly 2.0-12b: 17.4 (#296), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkDolly 2.0-12bLlama 13b
LMArena Non-English836819
LMArena Chinese836—

Instruction Following Dolly 2.0-12b leads

Dolly 2.0-12b: 38.7 (#304), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bLlama 13b
LMArena Instruction Following814781

Writing & Preference Dolly 2.0-12b leads

Dolly 2.0-12b: 15.2 (#311), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bLlama 13b
LMArena Text851834
LMArena Creative Writing864794
LMArena Multi-Turn740753

Frequently asked questions

Is Dolly 2.0-12b better than Llama 13b?

Dolly 2.0-12b is the stronger model overall, scoring 25.5 to 24.4 on the Noometry Index.

Is Dolly 2.0-12b or Llama 13b better for coding?

Dolly 2.0-12b scores higher on coding benchmarks: 23.4 versus 21.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Llama 13b share?

16 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper