Model comparison

GLM-5.1 vs phi-3-medium 14B

GLM-5.1 is the stronger model overall, scoring 47.8 to 29.7 on the Noometry Index.

Last verified . 2 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. GLM-5.1 scores higher in 3 categories and phi-3-medium 14B in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-5.1 leads 54.9 to 9.1.
  • The biggest single-benchmark swing is GPQA Diamond: 89.9% for GLM-5.1 and 27.6% for phi-3-medium 14B.

Side by side

GLM-5.1 and phi-3-medium 14B specifications
GLM-5.1phi-3-medium 14B
ProviderZ.ai (Zhipu)Microsoft
Noometry Index47.829.7
Released2026-04-072024-04-23
WeightsOpenOpen
Context window200K—
Max output131K—
Input $ / M tokens$1.40—
Output $ / M tokens$4.40—
Results tracked4113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.1 leads

GLM-5.1: 48.7 (#55), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
SWE-bench Verified74.2%—
LMArena WebDev1508—
SciCode43.8%—
WeirdML57.1%—
BigCodeBench Instruct—37.6%
LMArena Coding1485—
BigCodeBench Complete—48.7%
ALE-Bench887.1—

Agentic & Tool Use Not comparable

GLM-5.1: 24.9 (#113), phi-3-medium 14B: —

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
APEX-Agents40.9%—
ExploitBench18.1%—
GBAEval0%—
Vending-Bench 25,634—

Reasoning Not comparable

GLM-5.1: 39.1 (#60), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
Epoch Capabilities Index149.84121.23
SimpleBench55.1%—
NYT Connections (extended)77.7%—
CritPt4.6%—
Chess Puzzles19%—
Thematic Generalization69.8%—
LMArena Hard Prompts1472—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
HellaSwag—82.4%
WinoGrande—81.5%

Math GLM-5.1 leads

GLM-5.1: 49.7 (#60), phi-3-medium 14B: 27.3 (#250)

Knowledge GLM-5.1 leads

GLM-5.1: 54.9 (#50), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
GPQA Diamond89.9%27.6%
SimpleQA Verified34%—
LMArena Expert1476—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

GLM-5.1: 55.0 (#36), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
LMArena Non-English1447—
LMArena Chinese1515—
LMArena French1474—
LMArena German1465—
LMArena Japanese1434—
LMArena Korean1418—
LMArena Russian1454—
LMArena Spanish1469—

Instruction Following Not comparable

GLM-5.1: 76.3 (#42), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
LMArena Instruction Following1451—

Long Context Not comparable

GLM-5.1: 44.9 (#53), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
LMArena Longer Query1466—

Writing & Preference Not comparable

GLM-5.1: 66.9 (#31), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkGLM-5.1phi-3-medium 14B
LMArena Text1461—
LMArena Creative Writing1453—
EQ-Bench Creative Writing1592—
LMArena Multi-Turn1472—

Frequently asked questions

Is GLM-5.1 better than phi-3-medium 14B?

GLM-5.1 is the stronger model overall, scoring 47.8 to 29.7 on the Noometry Index.

Is GLM-5.1 or phi-3-medium 14B better for coding?

GLM-5.1 scores higher on coding benchmarks: 48.7 versus 36.8 in the Noometry coding category.

How many benchmarks do GLM-5.1 and phi-3-medium 14B share?

2 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper