Model comparison

Grok Build 0.1 vs Phi-4 Mini

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 30.9 on the Noometry Index. Phi-4 Mini costs 9.5× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • They share 2 benchmarks with published results for both. Grok Build 0.1 scores higher in 2 categories and Phi-4 Mini in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 28.1.
  • The biggest single-benchmark swing is SciCode: 50.2% for Grok Build 0.1 and 10.8% for Phi-4 Mini.
  • Phi-4 Mini is cheaper at $0.075 / $0.30 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 128K.
  • Phi-4 Mini has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Phi-4 Mini specifications
Grok Build 0.1Phi-4 Mini
ProviderxAIMicrosoft
Noometry Index36.430.9
Released2026-04-162024-12-11
WeightsProprietaryOpen
Context window256K128K
Max output256K4K
Input $ / M tokens$1$0.075
Output $ / M tokens$2$0.30
Results tracked33

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkGrok Build 0.1Phi-4 Mini
SciCode50.2%10.8%

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Phi-4 Mini: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Phi-4 Mini
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkGrok Build 0.1Phi-4 Mini
CritPt9.1%0%

Knowledge Not comparable

Grok Build 0.1: —, Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkGrok Build 0.1Phi-4 Mini
Vectara Hallucination Rate—23.5%

Frequently asked questions

Is Grok Build 0.1 better than Phi-4 Mini?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 30.9 on the Noometry Index. Phi-4 Mini costs 9.5× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Phi-4 Mini?

Phi-4 Mini is cheaper. It lists at $0.075 per million input tokens and $0.30 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Phi-4 Mini better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 28.1 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 128K.

How many benchmarks do Grok Build 0.1 and Phi-4 Mini share?

2 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper