Model comparison

Hy3 vs o1

Hy3 is the stronger model overall, scoring 44.2 to 40.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Hy3 scores higher in 5 categories and o1 in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hy3 leads 62.2 to 55.6.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $15 / $60 for o1.
  • Hy3 accepts more context: 262K tokens versus 200K.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Hy3 and o1 specifications
Hy3o1
ProviderTencentOpenAI
Noometry Index44.240.9
Released2026-07-062024-09-12
WeightsOpenProprietary
Context window262K200K
Max output128K100K
Input $ / M tokens$0.0825$15
Output $ / M tokens$0.33$60
Results tracked1952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Hy3: 46.8 (#63), o1: 46.1 (#70)

Coding benchmarks
BenchmarkHy3o1
LMArena Coding14641367
Aider Polyglot—61.7%
LMArena WebDev1508—
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Hy3: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkHy3o1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Hy3: 26.1 (#136), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkHy3o1
LMArena Hard Prompts14471371
SimpleBench—41.7%
NYT Connections (extended)41.2%—
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math Hy3 leads

Hy3: 40.1 (#93), o1: 36.1 (#175)

Math benchmarks
BenchmarkHy3o1
LMArena Math14751388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge Too close to call

Hy3: 40.8 (#114), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkHy3o1
LMArena Expert14601361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal Not comparable

Hy3: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkHy3o1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Hy3 leads

Hy3: 53.5 (#65), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkHy3o1
LMArena Non-English14261358
LMArena Chinese14931394
LMArena French14611344
LMArena German14391337
LMArena Japanese13921346
LMArena Korean13951396
LMArena Russian14321356
LMArena Spanish14561345

Instruction Following Too close to call

Hy3: 75.1 (#70), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkHy3o1
LMArena Instruction Following14261367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Hy3: 44.1 (#75), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkHy3o1
LMArena Longer Query14421378
Fiction.LiveBench—83.3%

Writing & Preference Hy3 leads

Hy3: 62.2 (#81), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkHy3o1
LMArena Text14391366
LMArena Creative Writing14021348
LMArena Multi-Turn14361369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Hy3 better than o1?

Hy3 is the stronger model overall, scoring 44.2 to 40.9 on the Noometry Index.

Which is cheaper, Hy3 or o1?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; o1 lists at $15 and $60.

Is Hy3 or o1 better for coding?

They score almost the same on coding (46.8 vs 46.1); test both on your own repository before choosing.

Which has the bigger context window?

Hy3 does, with 262K tokens against 200K.

How many benchmarks do Hy3 and o1 share?

17 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper