Model comparison

Dolly 2.0-12b vs Mistral

Mistral is the stronger model overall, scoring 29.9 to 25.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Dolly 2.0-12b scores higher in 1 category and Mistral in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral leads 37.0 to 15.2.
  • Dolly 2.0-12b has downloadable open weights; the other is API-only.

Side by side

Dolly 2.0-12b and Mistral specifications
Dolly 2.0-12bMistral
ProviderDatabricksMistral AI
Noometry Index25.529.9
Released2023-04-11—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1722

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Dolly 2.0-12b: 23.4 (#332), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Coding7761162

Reasoning Mistral leads

Dolly 2.0-12b: 15.3 (#316), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Hard Prompts8041149
Epoch Capabilities Index89.67—
HellaSwag70.8%—
PIQA75.4%—
WinoGrande61.8%—

Math Dolly 2.0-12b leads

Dolly 2.0-12b: 27.3 (#251), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Math8711180
Omni-MATH—7.2%

Knowledge Not comparable

Dolly 2.0-12b: —, Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkDolly 2.0-12bMistral
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
LMArena Expert—1125
ARC (AI2) Challenge39.6%—
BoolQ56.3%—
MMLU26.2%—
OpenBookQA39.2%—

Multilingual Mistral leads

Dolly 2.0-12b: 17.4 (#296), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Non-English8361129
LMArena Chinese8361109
LMArena French—1180
LMArena German—1155
LMArena Japanese—1013
LMArena Korean—1032
LMArena Russian—1168
LMArena Spanish—1143

Instruction Following Mistral leads

Dolly 2.0-12b: 38.7 (#304), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Instruction Following8141152
IFEval—56.8%

Long Context Not comparable

Dolly 2.0-12b: —, Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Longer Query—1153

Writing & Preference Mistral leads

Dolly 2.0-12b: 15.2 (#311), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bMistral
LMArena Text8511165
LMArena Creative Writing8641158
LMArena Multi-Turn7401147
WildBench—66%

Frequently asked questions

Is Dolly 2.0-12b better than Mistral?

Mistral is the stronger model overall, scoring 29.9 to 25.5 on the Noometry Index.

Is Dolly 2.0-12b or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Mistral share?

9 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper