Model comparison

Gemma 2B vs Phi 3 Mini 4k Instruct

Gemma 2B is the stronger model overall, scoring 29.6 to 27.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Phi 3 Mini 4k Instruct Microsoft

27.9

Rank #328 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemma 2B scores higher in 4 categories and Phi 3 Mini 4k Instruct in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 2B leads 18.8 to 14.1.

Side by side

Gemma 2B and Phi 3 Mini 4k Instruct specifications
Gemma 2BPhi 3 Mini 4k Instruct
ProviderGoogleMicrosoft
Noometry Index29.627.9
Released2024-02-212024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2335

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2B leads

Gemma 2B: 29.4 (#305), Phi 3 Mini 4k Instruct: 26.6 (#323)

Coding benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Coding10101093
HumanEval+20.7%59.1%
MBPP+34.1%54.2%
LiveBench Coding—15.5%

Reasoning Gemma 2B leads

Gemma 2B: 18.8 (#275), Phi 3 Mini 4k Instruct: 14.1 (#328)

Reasoning benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Hard Prompts9891072
BIG-Bench Hard35.2%71.7%
HellaSwag71.4%76.7%
WinoGrande65.4%70.8%
Chess Puzzles—0%
LiveBench Reasoning—26.8%
LiveBench Data Analysis—34.7%
Adversarial NLI—52.8%
Epoch Capabilities Index94.2—
LiveBench—22.4%
PIQA77.3%—

Math Gemma 2B leads

Gemma 2B: 30.0 (#239), Phi 3 Mini 4k Instruct: 26.6 (#257)

Math benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Math10091111
LiveBench Math—15.7%
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Phi 3 Mini 4k Instruct: 28.5 (#246)

Knowledge benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
ARC (AI2) Challenge42.1%84.9%
MMLU42.3%68.8%
TriviaQA53.2%64%
LMArena Expert—1045
BoolQ69.4%—
OpenBookQA—88%

Multilingual Phi 3 Mini 4k Instruct leads

Gemma 2B: 23.0 (#294), Phi 3 Mini 4k Instruct: 26.3 (#280)

Multilingual benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Non-English9581021
LMArena Chinese9861021
LMArena Russian9371022
LMArena French—1076
LMArena German—1044
LMArena Japanese—935
LMArena Korean—905
LMArena Spanish—1085

Instruction Following Too close to call

Gemma 2B: 48.5 (#302), Phi 3 Mini 4k Instruct: 47.7 (#303)

Instruction Following benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Instruction Following9701053
LiveBench Instruction Following—39.1%

Long Context Phi 3 Mini 4k Instruct leads

Gemma 2B: 29.9 (#291), Phi 3 Mini 4k Instruct: 31.7 (#276)

Long Context benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Longer Query9811044

Writing & Preference Phi 3 Mini 4k Instruct leads

Gemma 2B: 24.0 (#308), Phi 3 Mini 4k Instruct: 27.6 (#300)

Writing & Preference benchmarks
BenchmarkGemma 2BPhi 3 Mini 4k Instruct
LMArena Text10021073
LMArena Creative Writing9871037
LMArena Multi-Turn9451018
LiveBench Language—9.2%

Frequently asked questions

Is Gemma 2B better than Phi 3 Mini 4k Instruct?

Gemma 2B is the stronger model overall, scoring 29.6 to 27.9 on the Noometry Index.

Is Gemma 2B or Phi 3 Mini 4k Instruct better for coding?

Gemma 2B scores higher on coding benchmarks: 29.4 versus 26.6 in the Noometry coding category.

How many benchmarks do Gemma 2B and Phi 3 Mini 4k Instruct share?

19 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Phi 3 Mini 4k Instruct has 35.

Related comparisons

Go deeper