Databricks, open weights

Dolly 2.0-12b

Dolly 2.0-12b by Databricks ranks 342nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is math, where it ranks 251st.

Last verified

Specifications

Noometry rank
#342 of 354
Index score
25.5
Evidence
Confirmed 17 results
Provider
Databricks
Released
April 11, 2023
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Dolly 2.0-12b category scores
  1. Coding 23.4
  2. Reasoning 15.3
  3. Math 27.3
  4. Multilingual 17.4
  5. Instruction Following 38.7
  6. Writing & Preference 15.2
Dolly 2.0-12b category ranks
CategoryScoreRankResults
Coding23.4#3321
Reasoning15.3#3161
Math27.3#2511
Multilingual17.4#2961
Instruction Following38.7#3041
Writing & Preference15.2#3113

Strengths and weaknesses

Categories where Dolly 2.0-12b places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Dolly 2.0-12b: strongest categories
CategoryScorevs medianRank
Math27.3−9.3#251 of 327, top 77%
Reasoning15.3−8.3#316 of 350, top 91%
Coding23.4−15.3#332 of 340, top 98%

Weakest categories

Dolly 2.0-12b: weakest categories
CategoryScorevs medianRank
Writing & Preference15.2−38.6#311 of 312, top 100%
Instruction Following38.7−32.6#304 of 305, top 100%
Multilingual17.4−30.1#296 of 297, top 100%

Closest competitors

The models ranked just above and below Dolly 2.0-12b. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Dolly 2.0-12b
ModelRankScoreBlended $/MSpeed
Ministral 3B#33826.2$0.10—Compare
DeepSeek-R1-Distill-Qwen-1.5B#33926.1——Compare
Claude 3 Haiku#34025.9—41Compare
Gemma 2 9B#34125.9——Compare
GPT-4o mini#34325.5$0.26120Compare
Llama 3-8B#34425.5——Compare
Claude 2.1#34525.2——Compare
Claude 2#34625.0——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Dolly 2.0-12b Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding776#293 of 294, top 100%LMArena2026-10-08

Reasoning

Dolly 2.0-12b Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts804#296 of 297, top 100%LMArena2026-10-08
Epoch Capabilities Index89.67#210 of 213, top 99%Epoch AI2023-04-11
HellaSwag70.8%#27 of 29, top 94%Epoch AI
PIQA75.4%#27 of 27, top 100%Epoch AI
WinoGrande61.8%#38 of 43, top 89%Epoch AI

Math

Dolly 2.0-12b Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math871#284 of 285, top 100%LMArena2026-10-08

Knowledge

Dolly 2.0-12b Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
ARC (AI2) Challenge39.6%#35 of 39, top 90%Epoch AI
BoolQ56.3%#23 of 23, top 100%Epoch AI
MMLU26.2%#80 of 81, top 99%Epoch AI
OpenBookQA39.2%#18 of 19, top 95%Epoch AI

Multilingual

Dolly 2.0-12b Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English836#296 of 297, top 100%LMArena2026-10-08
LMArena Chinese836#285 of 285, top 100%LMArena2026-10-08

Instruction Following

Dolly 2.0-12b Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following814#297 of 298, top 100%LMArena2026-10-08

Writing & Preference

Dolly 2.0-12b Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text851#296 of 297, top 100%LMArena2026-10-08
LMArena Creative Writing864#294 of 295, top 100%LMArena2026-10-08
LMArena Multi-Turn740#295 of 295, top 100%LMArena2026-10-08

Compare Dolly 2.0-12b

Other Databricks models

Frequently asked questions

How good is Dolly 2.0-12b?

Dolly 2.0-12b by Databricks ranks 342nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is math, where it ranks 251st.

Is Dolly 2.0-12b open source?

Yes. Dolly 2.0-12b's weights are downloadable; check the license for commercial terms.

What are Dolly 2.0-12b's strengths and weaknesses?

Relative to other ranked models, Dolly 2.0-12b places best in math, reasoning, coding and lowest in writing & preference, instruction following, multilingual.

What is Dolly 2.0-12b best at?

Its best category is math, where it ranks 251st on Noometry.