Model comparison

Grok Build 0.1 vs Phi-4

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 31.2 on the Noometry Index. Phi-4 costs 14× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 17.7.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 128K.
  • Phi-4 has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Phi-4 specifications
Grok Build 0.1Phi-4
ProviderxAIMicrosoft
Noometry Index36.431.2
Released2026-04-162024-12-11
WeightsProprietaryOpen
Context window256K128K
Max output256K4K
Input $ / M tokens$1$0.07
Output $ / M tokens$2$0.14
Results tracked337

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkGrok Build 0.1Phi-4
SciCode50.2%—
BigCodeBench Instruct—45.5%
LiveBench Coding—30.7%
LMArena Coding—1231
BigCodeBench Complete—55.4%

Agentic & Tool Use Too close to call

Grok Build 0.1: 22.7 (#129), Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Phi-4
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkGrok Build 0.1Phi-4
CritPt9.1%—
Chess Puzzles—1%
LiveBench Reasoning—47.8%
LMArena Hard Prompts—1220
LiveBench Data Analysis—45.2%
Epoch Capabilities Index—130.42
LiveBench—41.6%

Math Not comparable

Grok Build 0.1: —, Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkGrok Build 0.1Phi-4
OTIS Mock AIME 2024-2025—13.8%
LiveBench Math—42%
LMArena Math—1246
MATH Level 5—64.9%

Knowledge Not comparable

Grok Build 0.1: —, Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkGrok Build 0.1Phi-4
GPQA Diamond—56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%
LMArena Expert—1203
MMLU—84.8%

Multilingual Not comparable

Grok Build 0.1: —, Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkGrok Build 0.1Phi-4
LMArena Non-English—1197
LMArena Chinese—1212
LMArena French—1224
LMArena German—1222
LMArena Japanese—1158
LMArena Korean—1151
LMArena Russian—1209
LMArena Spanish—1234

Instruction Following Not comparable

Grok Build 0.1: —, Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Phi-4
LiveBench Instruction Following—58.4%
LMArena Instruction Following—1201

Long Context Not comparable

Grok Build 0.1: —, Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkGrok Build 0.1Phi-4
LMArena Longer Query—1217

Writing & Preference Not comparable

Grok Build 0.1: —, Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Phi-4
LMArena Text—1217
LMArena Creative Writing—1182
Short-Story Creative Writing—62.6%
LMArena Multi-Turn—1206
LiveBench Language—25.6%

Frequently asked questions

Is Grok Build 0.1 better than Phi-4?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 31.2 on the Noometry Index. Phi-4 costs 14× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Phi-4 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 128K.

How many benchmarks do Grok Build 0.1 and Phi-4 share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper