Model comparison

Gemma 7B vs Phi 3 Small 8k Instruct

Gemma 7B and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (30.0 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 21 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemma 7B scores higher in 3 categories and Phi 3 Small 8k Instruct in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 7B leads 19.9 to 14.8.

Side by side

Gemma 7B and Phi 3 Small 8k Instruct specifications
Gemma 7BPhi 3 Small 8k Instruct
ProviderGoogleMicrosoft
Noometry Index30.029.3
Released2024-02-212024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2732

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 7B leads

Gemma 7B: 30.5 (#294), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Coding10481101
LiveBench Coding—20.3%
HumanEval+28.7%—
MBPP+43.4%—

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Hard Prompts10421100
Adversarial NLI48.7%58.1%
BIG-Bench Hard55.1%79.1%
HellaSwag82.2%77%
WinoGrande79%81.5%
LiveBench Reasoning—15.9%
LiveBench Data Analysis—30.3%
Epoch Capabilities Index111.99—
LiveBench—24%
PIQA81.2%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Math10661151
LiveBench Math—17.6%
GSM8K46.4%—

Knowledge Phi 3 Small 8k Instruct leads

Gemma 7B: 27.3 (#252), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Expert10011067
ARC (AI2) Challenge78.3%90.7%
MMLU66.1%75.7%
OpenBookQA78.6%88%
TriviaQA72.3%58.1%
BoolQ83.2%—

Multilingual Phi 3 Small 8k Instruct leads

Gemma 7B: 25.1 (#287), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Non-English9991058
LMArena Chinese10351061
LMArena French10251135
LMArena Russian9931111
LMArena German—1080
LMArena Japanese—966
LMArena Korean—894
LMArena Spanish—1111

Instruction Following Too close to call

Gemma 7B: 51.5 (#295), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Instruction Following10171087
LiveBench Instruction Following—47.2%

Long Context Phi 3 Small 8k Instruct leads

Gemma 7B: 31.1 (#282), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Longer Query10221088

Writing & Preference Phi 3 Small 8k Instruct leads

Gemma 7B: 27.1 (#302), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkGemma 7BPhi 3 Small 8k Instruct
LMArena Text10561110
LMArena Creative Writing10241083
LMArena Multi-Turn9631068
LiveBench Language—12.9%

Frequently asked questions

Is Gemma 7B better than Phi 3 Small 8k Instruct?

Gemma 7B and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (30.0 vs 29.3), so choose on price, context window or the category you care about most.

Is Gemma 7B or Phi 3 Small 8k Instruct better for coding?

Gemma 7B scores higher on coding benchmarks: 30.5 versus 27.9 in the Noometry coding category.

How many benchmarks do Gemma 7B and Phi 3 Small 8k Instruct share?

21 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper