Model comparison

Dolly 2.0-12b vs Mistral 7B

Dolly 2.0-12b is the stronger model overall, scoring 25.5 to 23.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Dolly 2.0-12b scores higher in 2 categories and Mistral 7B in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Dolly 2.0-12b leads 27.3 to 8.1.

Side by side

Dolly 2.0-12b and Mistral 7B specifications
Dolly 2.0-12bMistral 7B
ProviderDatabricksMistral AI
Noometry Index25.523.0
Released2023-04-112023-09-27
WeightsOpenOpen
Context window—8K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.25
Results tracked1737

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral 7B leads

Dolly 2.0-12b: 23.4 (#332), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Coding7761082
BigCodeBench Instruct—19.5%
BigCodeBench Complete—27.3%
HumanEval+—36%
MBPP+—42.1%

Reasoning Dolly 2.0-12b leads

Dolly 2.0-12b: 15.3 (#316), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Hard Prompts8041067
Epoch Capabilities Index89.67112.21
HellaSwag70.8%81%
PIQA75.4%83%
WinoGrande61.8%75.3%
Chess Puzzles—0%
DTBench—42.5%
Adversarial NLI—47.1%
BIG-Bench Hard—56.1%

Math Dolly 2.0-12b leads

Dolly 2.0-12b: 27.3 (#251), Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Math8711085
OTIS Mock AIME 2024-2025—0.3%
MATH Level 5—3.7%
GSM8K—54.4%

Knowledge Not comparable

Dolly 2.0-12b: —, Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkDolly 2.0-12bMistral 7B
ARC (AI2) Challenge39.6%78.6%
BoolQ56.3%87.4%
MMLU26.2%62.5%
OpenBookQA39.2%79.8%
GPQA Diamond—15.2%
LMArena Expert—1036
TriviaQA—75.2%

Multilingual Mistral 7B leads

Dolly 2.0-12b: 17.4 (#296), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Non-English8361012
LMArena Chinese8361009
LMArena French—1037
LMArena German—987
LMArena Japanese—878
LMArena Russian—1018
LMArena Spanish—1026

Instruction Following Mistral 7B leads

Dolly 2.0-12b: 38.7 (#304), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Instruction Following8141060

Long Context Not comparable

Dolly 2.0-12b: —, Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Longer Query—1060

Writing & Preference Mistral 7B leads

Dolly 2.0-12b: 15.2 (#311), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bMistral 7B
LMArena Text8511090
LMArena Creative Writing8641068
LMArena Multi-Turn7401062

Frequently asked questions

Is Dolly 2.0-12b better than Mistral 7B?

Dolly 2.0-12b is the stronger model overall, scoring 25.5 to 23.0 on the Noometry Index.

Is Dolly 2.0-12b or Mistral 7B better for coding?

Mistral 7B scores higher on coding benchmarks: 26.4 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Mistral 7B share?

17 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper