Model comparison

GPT-5.4 nano vs Hy3

Hy3 is the stronger model overall, scoring 44.2 to 41.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-5.4 nano scores higher in 2 categories and Hy3 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hy3 leads 62.2 to 55.7.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $0.20 / $1.25 for GPT-5.4 nano.
  • GPT-5.4 nano accepts more context: 400K tokens versus 262K.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 nano and Hy3 specifications
GPT-5.4 nanoHy3
ProviderOpenAITencent
Noometry Index41.944.2
Released2026-03-172026-07-06
WeightsProprietaryOpen
Context window400K262K
Max output128K128K
Input $ / M tokens$0.20$0.0825
Output $ / M tokens$1.25$0.33
Results tracked4019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

GPT-5.4 nano: 43.6 (#84), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Coding14051464
LMArena WebDev—1508
SciCode46.9%—
WeirdML49.2%—
ALE-Bench1,005—

Reasoning Hy3 leads

GPT-5.4 nano: 23.7 (#173), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Hard Prompts13811447
ARC-AGI-25.7%—
Kagi LLM Benchmark39.7%—
NYT Connections (extended)—41.2%
ARC-AGI-151.5%—
CritPt9.3%—
Chess Puzzles30%—
Mystery Game Puzzles9%—
DTBench80.3%—
LMCA36.9%—
Epoch Capabilities Index145.81—
ForecastBench57.3—

Math Too close to call

GPT-5.4 nano: 40.9 (#88), Hy3: 40.1 (#93)

Knowledge GPT-5.4 nano leads

GPT-5.4 nano: 41.9 (#103), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Expert13961460
GPQA Diamond78.5%—
SimpleQA Verified11.7%—
Vectara Hallucination Rate3.1%—

Multimodal Not comparable

GPT-5.4 nano: 36.7 (#78), Hy3: —

Multimodal benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Vision1196—

Multilingual Hy3 leads

GPT-5.4 nano: 48.6 (#140), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Non-English13591426
LMArena Chinese13921493
LMArena French13961461
LMArena German13671439
LMArena Japanese13431392
LMArena Korean13201395
LMArena Russian13631432
LMArena Spanish13711456

Instruction Following Hy3 leads

GPT-5.4 nano: 71.9 (#144), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Instruction Following13621426

Long Context Hy3 leads

GPT-5.4 nano: 41.6 (#137), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Longer Query13661442

Writing & Preference Hy3 leads

GPT-5.4 nano: 55.7 (#142), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkGPT-5.4 nanoHy3
LMArena Text13721439
LMArena Creative Writing13141402
LMArena Multi-Turn13821436

Frequently asked questions

Is GPT-5.4 nano better than Hy3?

Hy3 is the stronger model overall, scoring 44.2 to 41.9 on the Noometry Index.

Which is cheaper, GPT-5.4 nano or Hy3?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; GPT-5.4 nano lists at $0.20 and $1.25.

Is GPT-5.4 nano or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 43.6 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 nano does, with 400K tokens against 262K.

How many benchmarks do GPT-5.4 nano and Hy3 share?

17 benchmarks have published results for both models. GPT-5.4 nano has 40 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper