Model comparison

ERNIE 5.1 vs Phi-4 Mini

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 30.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in knowledge, where ERNIE 5.1 leads 41.9 to 25.3.
  • Phi-4 Mini has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Phi-4 Mini specifications
ERNIE 5.1Phi-4 Mini
ProviderBaiduMicrosoft
Noometry Index43.830.9
Released—2024-12-11
WeightsProprietaryOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.075
Output $ / M tokens—$0.30
Results tracked193

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
SciCode—10.8%
LMArena Coding1488—

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Phi-4 Mini: —

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
LMArena Search1227—

Reasoning Too close to call

ERNIE 5.1: 21.9 (#211), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
NYT Connections (extended)23.4%—
CritPt—0%
LMArena Hard Prompts1481—

Math Not comparable

ERNIE 5.1: 40.3 (#92), Phi-4 Mini: —

Math benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
LMArena Math1481—

Knowledge ERNIE 5.1 leads

ERNIE 5.1: 41.9 (#102), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
Vectara Hallucination Rate—23.5%
LMArena Expert1492—

Multilingual Not comparable

ERNIE 5.1: 55.5 (#29), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
LMArena Non-English1454—
LMArena Chinese1508—
LMArena French1488—
LMArena German1470—
LMArena Japanese1422—
LMArena Korean1427—
LMArena Russian1459—
LMArena Spanish1473—

Instruction Following Not comparable

ERNIE 5.1: 76.7 (#37), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
LMArena Instruction Following1460—

Long Context Not comparable

ERNIE 5.1: 44.7 (#59), Phi-4 Mini: —

Long Context benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
LMArena Longer Query1462—

Writing & Preference Not comparable

ERNIE 5.1: 65.1 (#52), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkERNIE 5.1Phi-4 Mini
LMArena Text1468—
LMArena Creative Writing1441—
LMArena Multi-Turn1471—

Frequently asked questions

Is ERNIE 5.1 better than Phi-4 Mini?

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 30.9 on the Noometry Index.

Is ERNIE 5.1 or Phi-4 Mini better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 28.1 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Phi-4 Mini share?

0 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper