Databricks, open weights

# Dolly 2.0-12b

> Dolly 2.0-12b by Databricks, released April 2023. Ranked #342 of 354 with a Noometry Index of 25.5. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/dolly-2-0-12b
- Last updated: 2026-10-10
- Title: Dolly 2.0-12b Benchmarks, Price & Rank (October 2026)

Dolly 2.0-12b by Databricks ranks 342nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is math, where it ranks 251st.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #342 of 354
- **Index score:** 25.5
- **Evidence:** Confirmed 17 results
- **Provider:** [![](/logos/databricks.svg) Databricks](https://noometry.com/providers/databricks)
- **Released:** April 11, 2023
- **Weights:** Open weights
- **Reasoning:** Unknown
- **Context window:** —
- **Max output:** —
- **Input price:** Not listed
- **Output price:** Not listed
- **Blended price:** Not listed
- **Output speed:** Not measured
- **Value:** Not ranked
- **Knowledge cutoff:** Unknown

## Category scores

Each category score combines every public result we have in that category.

Dolly 2.0-12b category scores

1.  Coding 23.4
2.  Reasoning 15.3
3.  Math 27.3
4.  Multilingual 17.4
5.  Instruction Following 38.7
6.  Writing & Preference 15.2
7.  010203040

Dolly 2.0-12b category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 23.4 | #332 | 1 |
| [Reasoning](https://noometry.com/best/reasoning) | 15.3 | #316 | 1 |
| [Math](https://noometry.com/best/math) | 27.3 | #251 | 1 |
| [Multilingual](https://noometry.com/best/multilingual) | 17.4 | #296 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 38.7 | #304 | 1 |
| [Writing & Preference](https://noometry.com/best/writing) | 15.2 | #311 | 3 |

## Strengths and weaknesses

Categories where Dolly 2.0-12b places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Dolly 2.0-12b: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Math](https://noometry.com/best/math) | 27.3 | −9.3 | #251 of 327, top 77% |
| [Reasoning](https://noometry.com/best/reasoning) | 15.3 | −8.3 | #316 of 350, top 91% |
| [Coding](https://noometry.com/best/coding) | 23.4 | −15.3 | #332 of 340, top 98% |

### Weakest categories

Dolly 2.0-12b: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Writing & Preference](https://noometry.com/best/writing) | 15.2 | −38.6 | #311 of 312, top 100% |
| [Instruction Following](https://noometry.com/best/instruction-following) | 38.7 | −32.6 | #304 of 305, top 100% |
| [Multilingual](https://noometry.com/best/multilingual) | 17.4 | −30.1 | #296 of 297, top 100% |

## Closest competitors

The models ranked just above and below Dolly 2.0-12b. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Dolly 2.0-12b
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Ministral 3B](https://noometry.com/models/ministral-3b) | #338 | 26.2 | $0.10 | — | [Compare](https://noometry.com/compare/dolly-2-0-12b-vs-ministral-3b) |
| [DeepSeek-R1-Distill-Qwen-1.5B](https://noometry.com/models/deepseek-r1-distill-qwen-1-5b) | #339 | 26.1 | — | — | [Compare](https://noometry.com/compare/deepseek-r1-distill-qwen-1-5b-vs-dolly-2-0-12b) |
| [Claude 3 Haiku](https://noometry.com/models/claude-3-haiku) | #340 | 25.9 | — | 41 | [Compare](https://noometry.com/compare/claude-3-haiku-vs-dolly-2-0-12b) |
| [Gemma 2 9B](https://noometry.com/models/gemma-2-9b) | #341 | 25.9 | — | — | [Compare](https://noometry.com/compare/dolly-2-0-12b-vs-gemma-2-9b) |
| [GPT-4o mini](https://noometry.com/models/gpt-4o-mini) | #343 | 25.5 | $0.26 | 120 | [Compare](https://noometry.com/compare/dolly-2-0-12b-vs-gpt-4o-mini) |
| [Llama 3-8B](https://noometry.com/models/llama-3-8b) | #344 | 25.5 | — | — | [Compare](https://noometry.com/compare/dolly-2-0-12b-vs-llama-3-8b) |
| [Claude 2.1](https://noometry.com/models/claude-2-1) | #345 | 25.2 | — | — | [Compare](https://noometry.com/compare/claude-2-1-vs-dolly-2-0-12b) |
| [Claude 2](https://noometry.com/models/claude-2) | #346 | 25.0 | — | — | [Compare](https://noometry.com/compare/claude-2-vs-dolly-2-0-12b) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Dolly 2.0-12b Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 776 | #293 of 294, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Reasoning

Dolly 2.0-12b Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 804 | #296 of 297, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 89.67 | #210 of 213, top 99% |  | [Epoch AI](https://epoch.ai/eci) | 2023-04-11 |
| [HellaSwag](https://noometry.com/benchmarks/hellaswag) | 70.8% | #27 of 29, top 94% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [PIQA](https://noometry.com/benchmarks/piqa) | 75.4% | #27 of 27, top 100% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WinoGrande](https://noometry.com/benchmarks/winogrande) | 61.8% | #38 of 43, top 89% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Math

Dolly 2.0-12b Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 871 | #284 of 285, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Knowledge

Dolly 2.0-12b Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [ARC (AI2) Challenge](https://noometry.com/benchmarks/arc-challenge) | 39.6% | #35 of 39, top 90% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BoolQ](https://noometry.com/benchmarks/boolq) | 56.3% | #23 of 23, top 100% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 26.2% | #80 of 81, top 99% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [OpenBookQA](https://noometry.com/benchmarks/openbookqa) | 39.2% | #18 of 19, top 95% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Multilingual

Dolly 2.0-12b Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 836 | #296 of 297, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 836 | #285 of 285, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

Dolly 2.0-12b Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 814 | #297 of 298, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

Dolly 2.0-12b Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 851 | #296 of 297, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 864 | #294 of 295, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 740 | #295 of 295, top 100% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## Compare Dolly 2.0-12b

-   [Dolly 2.0-12b vs Gemma 2 9B](https://noometry.com/compare/dolly-2-0-12b-vs-gemma-2-9b)
-   [Dolly 2.0-12b vs GPT-4o mini](https://noometry.com/compare/dolly-2-0-12b-vs-gpt-4o-mini)
-   [Dolly 2.0-12b vs Claude 3 Haiku](https://noometry.com/compare/claude-3-haiku-vs-dolly-2-0-12b)
-   [Dolly 2.0-12b vs Llama 3-8B](https://noometry.com/compare/dolly-2-0-12b-vs-llama-3-8b)
-   [Dolly 2.0-12b vs DeepSeek-R1-Distill-Qwen-1.5B](https://noometry.com/compare/deepseek-r1-distill-qwen-1-5b-vs-dolly-2-0-12b)
-   [Dolly 2.0-12b vs Claude 2.1](https://noometry.com/compare/claude-2-1-vs-dolly-2-0-12b)
-   [Dolly 2.0-12b vs GPT-6 Astra](https://noometry.com/compare/dolly-2-0-12b-vs-gpt-6-astra)
-   [Dolly 2.0-12b vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-dolly-2-0-12b)
-   [Dolly 2.0-12b vs Gemini 3.8 Flash](https://noometry.com/compare/dolly-2-0-12b-vs-gemini-3-8-flash)
-   [Dolly 2.0-12b vs Kimi K3](https://noometry.com/compare/dolly-2-0-12b-vs-kimi-k3)
-   [Dolly 2.0-12b vs Grok 4.6](https://noometry.com/compare/dolly-2-0-12b-vs-grok-4-6)
-   [Dolly 2.0-12b vs Qwen3.8 Max](https://noometry.com/compare/dolly-2-0-12b-vs-qwen3-8-max)
-   [Dolly 2.0-12b vs GLM-5.3](https://noometry.com/compare/dolly-2-0-12b-vs-glm-5-3)
-   [Dolly 2.0-12b vs Muse Spark 1.3](https://noometry.com/compare/dolly-2-0-12b-vs-muse-spark-1-3)

## Other Databricks models

-   [DBRX](https://noometry.com/models/dbrx)29.4

## Frequently asked questions

### How good is Dolly 2.0-12b?

Dolly 2.0-12b by Databricks ranks 342nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is math, where it ranks 251st.

### Is Dolly 2.0-12b open source?

Yes. Dolly 2.0-12b's weights are downloadable; check the license for commercial terms.

### What are Dolly 2.0-12b's strengths and weaknesses?

Relative to other ranked models, Dolly 2.0-12b places best in math, reasoning, coding and lowest in writing & preference, instruction following, multilingual.

### What is Dolly 2.0-12b best at?

Its best category is math, where it ranks 251st on Noometry.

### Cite this page

Noometry. (2026). Dolly 2.0-12b benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/dolly-2-0-12b

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/dolly-2-0-12b.md).
