Model comparison

Gemma 1.1 7b IT vs Phi-4

Gemma 1.1 7b IT and Phi-4 score almost the same on the Noometry Index (31.3 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Gemma 1.1 7b IT Google

31.3

Rank #277 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 1.1 7b IT scores higher in 2 categories and Phi-4 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 1.1 7b IT leads 32.0 to 20.8.

Side by side

Gemma 1.1 7b IT and Phi-4 specifications
Gemma 1.1 7b ITPhi-4
ProviderGoogleMicrosoft
Noometry Index31.331.2
Released—2024-12-11
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.07
Output $ / M tokens—$0.14
Results tracked1937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi-4 leads

Gemma 1.1 7b IT: 31.5 (#284), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Coding10841231
BigCodeBench Instruct—45.5%
LiveBench Coding—30.7%
BigCodeBench Complete—55.4%
HumanEval+35.4%—
MBPP+45%—

Agentic & Tool Use Not comparable

Gemma 1.1 7b IT: —, Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%

Reasoning Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 20.5 (#238), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Hard Prompts10711220
Chess Puzzles—1%
LiveBench Reasoning—47.8%
LiveBench Data Analysis—45.2%
Epoch Capabilities Index—130.42
LiveBench—41.6%

Math Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 32.0 (#220), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Math11071246
OTIS Mock AIME 2024-2025—13.8%
LiveBench Math—42%
MATH Level 5—64.9%

Knowledge Phi-4 leads

Gemma 1.1 7b IT: 28.3 (#247), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Expert10391203
GPQA Diamond—56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%
MMLU—84.8%

Multilingual Phi-4 leads

Gemma 1.1 7b IT: 28.1 (#273), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Non-English10521197
LMArena Chinese10611212
LMArena French10651224
LMArena German10541222
LMArena Japanese9711158
LMArena Korean9881151
LMArena Russian10461209
LMArena Spanish10491234

Instruction Following Phi-4 leads

Gemma 1.1 7b IT: 54.0 (#283), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Instruction Following10571201
LiveBench Instruction Following—58.4%

Long Context Phi-4 leads

Gemma 1.1 7b IT: 32.1 (#272), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Longer Query10561217

Writing & Preference Phi-4 leads

Gemma 1.1 7b IT: 30.4 (#288), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkGemma 1.1 7b ITPhi-4
LMArena Text10941217
LMArena Creative Writing10601182
LMArena Multi-Turn10401206
Short-Story Creative Writing—62.6%
LiveBench Language—25.6%

Frequently asked questions

Is Gemma 1.1 7b IT better than Phi-4?

Gemma 1.1 7b IT and Phi-4 score almost the same on the Noometry Index (31.3 vs 31.2), so choose on price, context window or the category you care about most.

Is Gemma 1.1 7b IT or Phi-4 better for coding?

Phi-4 scores higher on coding benchmarks: 34.4 versus 31.5 in the Noometry coding category.

How many benchmarks do Gemma 1.1 7b IT and Phi-4 share?

17 benchmarks have published results for both models. Gemma 1.1 7b IT has 19 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper