Model comparison

Dolly 2.0-12b vs GPT-4

GPT-4 is the stronger model overall, scoring 29.1 to 25.5 on the Noometry Index.

Last verified . 13 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Dolly 2.0-12b scores higher in 1 category and GPT-4 in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where GPT-4 leads 65.3 to 38.7.
  • Dolly 2.0-12b has downloadable open weights; the other is API-only.

Side by side

Dolly 2.0-12b and GPT-4 specifications
Dolly 2.0-12bGPT-4
ProviderDatabricksOpenAI
Noometry Index25.529.1
Released2023-04-112023-03-14
WeightsOpenProprietary
Context window—8K
Max output—8K
Input $ / M tokens—$30
Output $ / M tokens—$60
Results tracked1738

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

Dolly 2.0-12b: 23.4 (#332), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Coding7761254
WeirdML—12.4%
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
HumanEval+—79.3%

Agentic & Tool Use Not comparable

Dolly 2.0-12b: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkDolly 2.0-12bGPT-4
METR Time Horizons—36.1%

Reasoning GPT-4 leads

Dolly 2.0-12b: 15.3 (#316), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Hard Prompts8041241
Epoch Capabilities Index89.67125.89
HellaSwag70.8%95.3%
WinoGrande61.8%87.5%
Chess Puzzles—4%
Mystery Game Puzzles—12%
DTBench—62.7%
LMCA—17.1%
BIG-Bench Hard—75.1%
ForecastBench—57.8
PIQA75.4%—

Math Dolly 2.0-12b leads

Dolly 2.0-12b: 27.3 (#251), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Math8711269
OTIS Mock AIME 2024-2025—1.1%
MATH Level 5—23%
GSM8K—92%

Knowledge Not comparable

Dolly 2.0-12b: —, GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkDolly 2.0-12bGPT-4
MMLU26.2%86.4%
GPQA Diamond—35.7%
LMArena Expert—1211
ARC (AI2) Challenge39.6%—
BoolQ56.3%—
OpenBookQA39.2%—
TriviaQA—84.8%

Multilingual GPT-4 leads

Dolly 2.0-12b: 17.4 (#296), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Non-English8361246
LMArena Chinese8361242
LMArena French—1283
LMArena German—1251
LMArena Japanese—1209
LMArena Korean—1184
LMArena Russian—1251
LMArena Spanish—1261

Instruction Following GPT-4 leads

Dolly 2.0-12b: 38.7 (#304), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Instruction Following8141241

Long Context Not comparable

Dolly 2.0-12b: —, GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Longer Query—1244

Writing & Preference GPT-4 leads

Dolly 2.0-12b: 15.2 (#311), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bGPT-4
LMArena Text8511263
LMArena Creative Writing8641244
LMArena Multi-Turn7401257
EQ-Bench Creative Writing—752

Frequently asked questions

Is Dolly 2.0-12b better than GPT-4?

GPT-4 is the stronger model overall, scoring 29.1 to 25.5 on the Noometry Index.

Is Dolly 2.0-12b or GPT-4 better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and GPT-4 share?

13 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper