Model comparison

Codellama 34b Instruct vs Dolly 2.0-12b

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 25.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Codellama 34b Instruct scores higher in 6 categories and Dolly 2.0-12b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Codellama 34b Instruct leads 52.2 to 38.7.

Side by side

Codellama 34b Instruct and Dolly 2.0-12b specifications
Codellama 34b InstructDolly 2.0-12b
ProviderMetaDatabricks
Noometry Index30.825.5
Released—2023-04-11
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 34b Instruct leads

Codellama 34b Instruct: 28.5 (#314), Dolly 2.0-12b: 23.4 (#332)

Coding benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Coding1046776
BigCodeBench Instruct29%—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Reasoning Codellama 34b Instruct leads

Codellama 34b Instruct: 19.6 (#255), Dolly 2.0-12b: 15.3 (#316)

Reasoning benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Hard Prompts1032804
Epoch Capabilities Index—89.67
HellaSwag—70.8%
PIQA—75.4%
WinoGrande—61.8%

Math Codellama 34b Instruct leads

Codellama 34b Instruct: 31.0 (#230), Dolly 2.0-12b: 27.3 (#251)

Math benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Math1056871

Knowledge Not comparable

Codellama 34b Instruct: —, Dolly 2.0-12b: —

Knowledge benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
ARC (AI2) Challenge—39.6%
BoolQ—56.3%
MMLU—26.2%
OpenBookQA—39.2%

Multilingual Codellama 34b Instruct leads

Codellama 34b Instruct: 25.8 (#284), Dolly 2.0-12b: 17.4 (#296)

Multilingual benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Non-English1011836
LMArena Chinese976836

Instruction Following Codellama 34b Instruct leads

Codellama 34b Instruct: 52.2 (#291), Dolly 2.0-12b: 38.7 (#304)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Instruction Following1028814

Long Context Not comparable

Codellama 34b Instruct: 30.9 (#284), Dolly 2.0-12b: —

Long Context benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Longer Query1013—

Writing & Preference Codellama 34b Instruct leads

Codellama 34b Instruct: 28.2 (#297), Dolly 2.0-12b: 15.2 (#311)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructDolly 2.0-12b
LMArena Text1066851
LMArena Creative Writing1032864
LMArena Multi-Turn1015740

Frequently asked questions

Is Codellama 34b Instruct better than Dolly 2.0-12b?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 25.5 on the Noometry Index.

Is Codellama 34b Instruct or Dolly 2.0-12b better for coding?

Codellama 34b Instruct scores higher on coding benchmarks: 28.5 versus 23.4 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Dolly 2.0-12b share?

9 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Dolly 2.0-12b has 17.

Related comparisons

Go deeper