Databricks, open weights
Dolly 2.0-12b
Dolly 2.0-12b by Databricks ranks 342nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is math, where it ranks 251st.
Last verified
Specifications
- Noometry rank
- #342 of 354
- Index score
- 25.5
- Evidence
- Confirmed 17 results
- Provider
Databricks
- Released
- April 11, 2023
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 23.4
- Reasoning 15.3
- Math 27.3
- Multilingual 17.4
- Instruction Following 38.7
- Writing & Preference 15.2
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 23.4 | #332 | 1 |
| Reasoning | 15.3 | #316 | 1 |
| Math | 27.3 | #251 | 1 |
| Multilingual | 17.4 | #296 | 1 |
| Instruction Following | 38.7 | #304 | 1 |
| Writing & Preference | 15.2 | #311 | 3 |
Strengths and weaknesses
Categories where Dolly 2.0-12b places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 15.2 | −38.6 | #311 of 312, top 100% |
| Instruction Following | 38.7 | −32.6 | #304 of 305, top 100% |
| Multilingual | 17.4 | −30.1 | #296 of 297, top 100% |
Closest competitors
The models ranked just above and below Dolly 2.0-12b. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Ministral 3B | #338 | 26.2 | $0.10 | — | Compare |
| DeepSeek-R1-Distill-Qwen-1.5B | #339 | 26.1 | — | — | Compare |
| Claude 3 Haiku | #340 | 25.9 | — | 41 | Compare |
| Gemma 2 9B | #341 | 25.9 | — | — | Compare |
| GPT-4o mini | #343 | 25.5 | $0.26 | 120 | Compare |
| Llama 3-8B | #344 | 25.5 | — | — | Compare |
| Claude 2.1 | #345 | 25.2 | — | — | Compare |
| Claude 2 | #346 | 25.0 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 776 | #293 of 294, top 100% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 804 | #296 of 297, top 100% | LMArena | 2026-10-08 | |
| Epoch Capabilities Index | 89.67 | #210 of 213, top 99% | Epoch AI | 2023-04-11 | |
| HellaSwag | 70.8% | #27 of 29, top 94% | Epoch AI | ||
| PIQA | 75.4% | #27 of 27, top 100% | Epoch AI | ||
| WinoGrande | 61.8% | #38 of 43, top 89% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 871 | #284 of 285, top 100% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC (AI2) Challenge | 39.6% | #35 of 39, top 90% | Epoch AI | ||
| BoolQ | 56.3% | #23 of 23, top 100% | Epoch AI | ||
| MMLU | 26.2% | #80 of 81, top 99% | Epoch AI | ||
| OpenBookQA | 39.2% | #18 of 19, top 95% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 836 | #296 of 297, top 100% | LMArena | 2026-10-08 | |
| LMArena Chinese | 836 | #285 of 285, top 100% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 814 | #297 of 298, top 100% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 851 | #296 of 297, top 100% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 864 | #294 of 295, top 100% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 740 | #295 of 295, top 100% | LMArena | 2026-10-08 |
Compare Dolly 2.0-12b
- Dolly 2.0-12b vs Gemma 2 9B
- Dolly 2.0-12b vs GPT-4o mini
- Dolly 2.0-12b vs Claude 3 Haiku
- Dolly 2.0-12b vs Llama 3-8B
- Dolly 2.0-12b vs DeepSeek-R1-Distill-Qwen-1.5B
- Dolly 2.0-12b vs Claude 2.1
- Dolly 2.0-12b vs GPT-6 Astra
- Dolly 2.0-12b vs Claude Fable 5.1
- Dolly 2.0-12b vs Gemini 3.8 Flash
- Dolly 2.0-12b vs Kimi K3
- Dolly 2.0-12b vs Grok 4.6
- Dolly 2.0-12b vs Qwen3.8 Max
- Dolly 2.0-12b vs GLM-5.3
- Dolly 2.0-12b vs Muse Spark 1.3
Other Databricks models
- DBRX29.4
Frequently asked questions
How good is Dolly 2.0-12b?
Dolly 2.0-12b by Databricks ranks 342nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is math, where it ranks 251st.
Is Dolly 2.0-12b open source?
Yes. Dolly 2.0-12b's weights are downloadable; check the license for commercial terms.
What are Dolly 2.0-12b's strengths and weaknesses?
Relative to other ranked models, Dolly 2.0-12b places best in math, reasoning, coding and lowest in writing & preference, instruction following, multilingual.
What is Dolly 2.0-12b best at?
Its best category is math, where it ranks 251st on Noometry.