Model comparison

Gemma 7B vs Phi 3 Mini 4k Instruct

Gemma 7B is the stronger model overall, scoring 30.0 to 27.9 on the Noometry Index.

Last verified . 23 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Phi 3 Mini 4k Instruct Microsoft

27.9

Rank #328 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Gemma 7B scores higher in 4 categories and Phi 3 Mini 4k Instruct in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 7B leads 19.9 to 14.1.

Side by side

Gemma 7B and Phi 3 Mini 4k Instruct specifications
Gemma 7BPhi 3 Mini 4k Instruct
ProviderGoogleMicrosoft
Noometry Index30.027.9
Released2024-02-212024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2735

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 7B leads

Gemma 7B: 30.5 (#294), Phi 3 Mini 4k Instruct: 26.6 (#323)

Coding benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Coding10481093
HumanEval+28.7%59.1%
MBPP+43.4%54.2%
LiveBench Coding—15.5%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Phi 3 Mini 4k Instruct: 14.1 (#328)

Reasoning benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Hard Prompts10421072
Adversarial NLI48.7%52.8%
BIG-Bench Hard55.1%71.7%
HellaSwag82.2%76.7%
WinoGrande79%70.8%
Chess Puzzles—0%
LiveBench Reasoning—26.8%
LiveBench Data Analysis—34.7%
Epoch Capabilities Index111.99—
LiveBench—22.4%
PIQA81.2%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Phi 3 Mini 4k Instruct: 26.6 (#257)

Math benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Math10661111
LiveBench Math—15.7%
GSM8K46.4%—

Knowledge Phi 3 Mini 4k Instruct leads

Gemma 7B: 27.3 (#252), Phi 3 Mini 4k Instruct: 28.5 (#246)

Knowledge benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Expert10011045
ARC (AI2) Challenge78.3%84.9%
MMLU66.1%68.8%
OpenBookQA78.6%88%
TriviaQA72.3%64%
BoolQ83.2%—

Multilingual Phi 3 Mini 4k Instruct leads

Gemma 7B: 25.1 (#287), Phi 3 Mini 4k Instruct: 26.3 (#280)

Multilingual benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Non-English9991021
LMArena Chinese10351021
LMArena French10251076
LMArena Russian9931022
LMArena German—1044
LMArena Japanese—935
LMArena Korean—905
LMArena Spanish—1085

Instruction Following Gemma 7B leads

Gemma 7B: 51.5 (#295), Phi 3 Mini 4k Instruct: 47.7 (#303)

Instruction Following benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Instruction Following10171053
LiveBench Instruction Following—39.1%

Long Context Too close to call

Gemma 7B: 31.1 (#282), Phi 3 Mini 4k Instruct: 31.7 (#276)

Long Context benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Longer Query10221044

Writing & Preference Too close to call

Gemma 7B: 27.1 (#302), Phi 3 Mini 4k Instruct: 27.6 (#300)

Writing & Preference benchmarks
BenchmarkGemma 7BPhi 3 Mini 4k Instruct
LMArena Text10561073
LMArena Creative Writing10241037
LMArena Multi-Turn9631018
LiveBench Language—9.2%

Frequently asked questions

Is Gemma 7B better than Phi 3 Mini 4k Instruct?

Gemma 7B is the stronger model overall, scoring 30.0 to 27.9 on the Noometry Index.

Is Gemma 7B or Phi 3 Mini 4k Instruct better for coding?

Gemma 7B scores higher on coding benchmarks: 30.5 versus 26.6 in the Noometry coding category.

How many benchmarks do Gemma 7B and Phi 3 Mini 4k Instruct share?

23 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Phi 3 Mini 4k Instruct has 35.

Related comparisons

Go deeper