Model comparison

Gemini 1.5 Flash (May 2024) vs Phi 3 Mini 4k Instruct June 2024

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.3 on the Noometry Index.

Last verified . 15 shared benchmarks.

Summary

  • They share 15 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 6 categories and Phi 3 Mini 4k Instruct June 2024 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash (May 2024) leads 48.7 to 29.6.
  • Phi 3 Mini 4k Instruct June 2024 has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Phi 3 Mini 4k Instruct June 2024 specifications
Gemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
ProviderGoogleMicrosoft
Noometry Index33.231.3
Released2024-05-14—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), Phi 3 Mini 4k Instruct June 2024: 31.8 (#279)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Coding12611093
WeirdML24.9%—
BigCodeBench Instruct43.5%—
BigCodeBench Complete55.1%—
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Not comparable

Gemini 1.5 Flash (May 2024): 26.6 (#102), Phi 3 Mini 4k Instruct June 2024: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
BALROG14.6%—

Reasoning Too close to call

Gemini 1.5 Flash (May 2024): 21.7 (#215), Phi 3 Mini 4k Instruct June 2024: 20.8 (#231)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Hard Prompts12571087
DTBench53.8%—
Epoch Capabilities Index129.36—
ForecastBench53.9—
PIQA87.5%—

Math Phi 3 Mini 4k Instruct June 2024 leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Phi 3 Mini 4k Instruct June 2024: 33.0 (#208)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Math12691152
OTIS Mock AIME 2024-202516.3%—
Omni-MATH30.4%—
MATH Level 561.9%—
FrontierMath (Feb 2025 set)0%—
GSM8K82.4%—

Knowledge Phi 3 Mini 4k Instruct June 2024 leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Phi 3 Mini 4k Instruct June 2024: 28.6 (#244)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Expert12331051
GPQA Diamond47.3%—
MMLU-Pro67.8%—
GPQA (HELM)43.7%—
BoolQ85.8%—
MMLU77.9%—

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), Phi 3 Mini 4k Instruct June 2024: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—

Multilingual Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Phi 3 Mini 4k Instruct June 2024: 25.9 (#282)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Non-English12781013
LMArena Chinese12951033
LMArena German12621031
LMArena Japanese1252954
LMArena Korean1221880
LMArena Russian12881019
LMArena French1258—
LMArena Spanish1243—

Instruction Following Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Phi 3 Mini 4k Instruct June 2024: 54.1 (#282)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Instruction Following12581058
IFEval83.1%—

Long Context Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), Phi 3 Mini 4k Instruct June 2024: 31.7 (#277)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Longer Query12841042

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Phi 3 Mini 4k Instruct June 2024: 29.6 (#294)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Phi 3 Mini 4k Instruct June 2024
LMArena Text12871080
LMArena Creative Writing12851045
LMArena Multi-Turn12531049
WildBench79.2%—

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Phi 3 Mini 4k Instruct June 2024?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.3 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Phi 3 Mini 4k Instruct June 2024 better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 31.8 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Phi 3 Mini 4k Instruct June 2024 share?

15 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Phi 3 Mini 4k Instruct June 2024 has 15.

Related comparisons

Go deeper