Model comparison

Dolly 2.0-12b vs phi-3-medium 14B

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 25.5 on the Noometry Index.

Last verified . 6 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 6 benchmarks with published results for both. Dolly 2.0-12b scores higher in 0 categories and phi-3-medium 14B in 2 categories; one gap is clear of the uncertainty.
  • The widest gap is in coding, where phi-3-medium 14B leads 36.8 to 23.4.

Side by side

Dolly 2.0-12b and phi-3-medium 14B specifications
Dolly 2.0-12bphi-3-medium 14B
ProviderDatabricksMicrosoft
Noometry Index25.529.7
Released2023-04-112024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

Dolly 2.0-12b: 23.4 (#332), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
BigCodeBench Instruct—37.6%
LMArena Coding776—
BigCodeBench Complete—48.7%

Reasoning Not comparable

Dolly 2.0-12b: 15.3 (#316), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
Epoch Capabilities Index89.67121.23
HellaSwag70.8%82.4%
WinoGrande61.8%81.5%
LMArena Hard Prompts804—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
PIQA75.4%—

Math Too close to call

Dolly 2.0-12b: 27.3 (#251), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
LMArena Math871—
MATH Level 5—17.6%

Knowledge Not comparable

Dolly 2.0-12b: —, phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
ARC (AI2) Challenge39.6%91.6%
MMLU26.2%78%
OpenBookQA39.2%87.4%
GPQA Diamond—27.6%
BoolQ56.3%—
TriviaQA—73.9%

Multilingual Not comparable

Dolly 2.0-12b: 17.4 (#296), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
LMArena Non-English836—
LMArena Chinese836—

Instruction Following Not comparable

Dolly 2.0-12b: 38.7 (#304), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
LMArena Instruction Following814—

Writing & Preference Not comparable

Dolly 2.0-12b: 15.2 (#311), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bphi-3-medium 14B
LMArena Text851—
LMArena Creative Writing864—
LMArena Multi-Turn740—

Frequently asked questions

Is Dolly 2.0-12b better than phi-3-medium 14B?

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 25.5 on the Noometry Index.

Is Dolly 2.0-12b or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and phi-3-medium 14B share?

6 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper