Model comparison

Dolly 2.0-12b vs Gemma 3n E4b IT

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 25.5 on the Noometry Index.

Last verified . 9 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Dolly 2.0-12b scores higher in 0 categories and Gemma 3n E4b IT in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemma 3n E4b IT leads 50.1 to 15.2.

Side by side

Dolly 2.0-12b and Gemma 3n E4b IT specifications
Dolly 2.0-12bGemma 3n E4b IT
ProviderDatabricksGoogle
Noometry Index25.537.3
Released2023-04-11—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1718

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3n E4b IT leads

Dolly 2.0-12b: 23.4 (#332), Gemma 3n E4b IT: 37.0 (#198)

Coding benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Coding7761268

Reasoning Gemma 3n E4b IT leads

Dolly 2.0-12b: 15.3 (#316), Gemma 3n E4b IT: 19.9 (#247)

Reasoning benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Hard Prompts8041284
Kagi LLM Benchmark—31.5%
Epoch Capabilities Index89.67—
HellaSwag70.8%—
PIQA75.4%—
WinoGrande61.8%—

Math Gemma 3n E4b IT leads

Dolly 2.0-12b: 27.3 (#251), Gemma 3n E4b IT: 35.1 (#188)

Math benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Math8711251

Knowledge Not comparable

Dolly 2.0-12b: —, Gemma 3n E4b IT: 34.2 (#198)

Knowledge benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Expert—1246
ARC (AI2) Challenge39.6%—
BoolQ56.3%—
MMLU26.2%—
OpenBookQA39.2%—

Multilingual Gemma 3n E4b IT leads

Dolly 2.0-12b: 17.4 (#296), Gemma 3n E4b IT: 43.4 (#183)

Multilingual benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Non-English8361285
LMArena Chinese8361309
LMArena French—1330
LMArena German—1311
LMArena Japanese—1272
LMArena Korean—1259
LMArena Russian—1288
LMArena Spanish—1305

Instruction Following Gemma 3n E4b IT leads

Dolly 2.0-12b: 38.7 (#304), Gemma 3n E4b IT: 66.1 (#210)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Instruction Following8141255

Long Context Not comparable

Dolly 2.0-12b: —, Gemma 3n E4b IT: 38.7 (#191)

Long Context benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Longer Query—1276

Writing & Preference Gemma 3n E4b IT leads

Dolly 2.0-12b: 15.2 (#311), Gemma 3n E4b IT: 50.1 (#186)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bGemma 3n E4b IT
LMArena Text8511306
LMArena Creative Writing8641287
LMArena Multi-Turn7401276

Frequently asked questions

Is Dolly 2.0-12b better than Gemma 3n E4b IT?

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 25.5 on the Noometry Index.

Is Dolly 2.0-12b or Gemma 3n E4b IT better for coding?

Gemma 3n E4b IT scores higher on coding benchmarks: 37.0 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Gemma 3n E4b IT share?

9 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Gemma 3n E4b IT has 18.

Related comparisons

Go deeper