Model comparison

GPT-5.5 Instant vs Hy3

Hy3 is the stronger model overall, scoring 44.2 to 42.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-5.5 Instant scores higher in 1 category and Hy3 in 7 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy3 leads 40.1 to 26.5.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

GPT-5.5 Instant and Hy3 specifications
GPT-5.5 InstantHy3
ProviderOpenAITencent
Noometry Index42.744.2
Released2026-05-052026-07-06
WeightsProprietaryOpen
Context window—262K
Max output—128K
Input $ / M tokens—$0.0825
Output $ / M tokens—$0.33
Results tracked2719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

GPT-5.5 Instant: 44.3 (#74), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Coding14331464
LMArena WebDev—1508
SciCode48.6%—

Reasoning Hy3 leads

GPT-5.5 Instant: 24.9 (#155), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Hard Prompts14261447
NYT Connections (extended)—41.2%
CritPt0%—
Chess Puzzles12%—
Epoch Capabilities Index142.52—

Math Hy3 leads

GPT-5.5 Instant: 26.5 (#259), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Math14201475
FrontierMath (Tiers 1-3)26.3%—
FrontierMath Tier 42.4%—
OTIS Mock AIME 2024-202568.1%—

Knowledge GPT-5.5 Instant leads

GPT-5.5 Instant: 48.9 (#74), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Expert14091460
GPQA Diamond82.5%—

Multimodal Not comparable

GPT-5.5 Instant: 40.0 (#52), Hy3: —

Multimodal benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Vision1250—
LMArena Document1403—

Multilingual Too close to call

GPT-5.5 Instant: 52.8 (#80), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Non-English14171426
LMArena Chinese14561493
LMArena French14281461
LMArena German14111439
LMArena Japanese14081392
LMArena Korean13921395
LMArena Russian14311432
LMArena Spanish14291456

Instruction Following Too close to call

GPT-5.5 Instant: 74.2 (#100), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Instruction Following14061426

Long Context Too close to call

GPT-5.5 Instant: 43.4 (#96), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Longer Query14221442

Writing & Preference Too close to call

GPT-5.5 Instant: 61.8 (#85), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkGPT-5.5 InstantHy3
LMArena Text14191439
LMArena Creative Writing14191402
LMArena Multi-Turn14331436

Frequently asked questions

Is GPT-5.5 Instant better than Hy3?

Hy3 is the stronger model overall, scoring 44.2 to 42.7 on the Noometry Index.

Is GPT-5.5 Instant or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 44.3 in the Noometry coding category.

How many benchmarks do GPT-5.5 Instant and Hy3 share?

17 benchmarks have published results for both models. GPT-5.5 Instant has 27 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper