Model comparison

Hy3 vs Phi-4 Mini

Hy3 is the stronger model overall, scoring 44.2 to 30.9 on the Noometry Index. Phi-4 Mini costs 1.8× less per token, which makes it the better buy when Hy3's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in coding, where Hy3 leads 46.8 to 28.1.
  • Phi-4 Mini is cheaper at $0.075 / $0.30 per million input/output tokens, against $0.13 / $0.53 for Hy3.
  • Hy3 accepts more context: 262K tokens versus 128K.

Side by side

Hy3 and Phi-4 Mini specifications
Hy3Phi-4 Mini
ProviderTencentMicrosoft
Noometry Index44.230.9
Released2026-07-062024-12-11
WeightsOpenOpen
Context window262K128K
Max output128K4K
Input $ / M tokens$0.13$0.075
Output $ / M tokens$0.53$0.30
Results tracked193

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Hy3: 46.8 (#63), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkHy3Phi-4 Mini
LMArena WebDev1508—
SciCode—10.8%
LMArena Coding1464—

Reasoning Hy3 leads

Hy3: 26.1 (#136), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkHy3Phi-4 Mini
NYT Connections (extended)41.2%—
CritPt—0%
LMArena Hard Prompts1447—

Math Not comparable

Hy3: 40.1 (#93), Phi-4 Mini: —

Math benchmarks
BenchmarkHy3Phi-4 Mini
LMArena Math1475—

Knowledge Hy3 leads

Hy3: 40.8 (#114), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkHy3Phi-4 Mini
Vectara Hallucination Rate—23.5%
LMArena Expert1460—

Multilingual Not comparable

Hy3: 53.5 (#65), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkHy3Phi-4 Mini
LMArena Non-English1426—
LMArena Chinese1493—
LMArena French1461—
LMArena German1439—
LMArena Japanese1392—
LMArena Korean1395—
LMArena Russian1432—
LMArena Spanish1456—

Instruction Following Not comparable

Hy3: 75.1 (#70), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkHy3Phi-4 Mini
LMArena Instruction Following1426—

Long Context Not comparable

Hy3: 44.1 (#75), Phi-4 Mini: —

Long Context benchmarks
BenchmarkHy3Phi-4 Mini
LMArena Longer Query1442—

Writing & Preference Not comparable

Hy3: 62.2 (#81), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkHy3Phi-4 Mini
LMArena Text1439—
LMArena Creative Writing1402—
LMArena Multi-Turn1436—

Frequently asked questions

Is Hy3 better than Phi-4 Mini?

Hy3 is the stronger model overall, scoring 44.2 to 30.9 on the Noometry Index. Phi-4 Mini costs 1.8× less per token, which makes it the better buy when Hy3's lead doesn't matter for your workload.

Which is cheaper, Hy3 or Phi-4 Mini?

Phi-4 Mini is cheaper. It lists at $0.075 per million input tokens and $0.30 per million output tokens; Hy3 lists at $0.13 and $0.53.

Is Hy3 or Phi-4 Mini better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 28.1 in the Noometry coding category.

Which has the bigger context window?

Hy3 does, with 262K tokens against 128K.

How many benchmarks do Hy3 and Phi-4 Mini share?

0 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper