Model comparison

Claude 3.5 Haiku vs Hy3

Hy3 is the stronger model overall, scoring 44.2 to 29.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Hy3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy3 leads 40.1 to 14.7.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Hy3 specifications
Claude 3.5 HaikuHy3
ProviderAnthropicTencent
Noometry Index29.244.2
Released2024-10-222026-07-06
WeightsProprietaryOpen
Context window—262K
Max output—128K
Input $ / M tokens—$0.0825
Output $ / M tokens—$0.33
Results tracked4919

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Claude 3.5 Haiku: 32.9 (#265), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Coding12861464
Aider Polyglot28%—
LMArena WebDev—1508
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Hy3: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuHy3
BALROG19.3%—

Reasoning Hy3 leads

Claude 3.5 Haiku: 17.7 (#290), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Hard Prompts12511447
NYT Connections (extended)—41.2%
CritPt0%—
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Hy3 leads

Claude 3.5 Haiku: 14.7 (#300), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Math12441475
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Hy3 leads

Claude 3.5 Haiku: 18.7 (#281), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Expert12081460
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Hy3: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Vision1092—
GeoBench34%—

Multilingual Hy3 leads

Claude 3.5 Haiku: 40.0 (#218), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Non-English12381426
LMArena Chinese12291493
LMArena French12641461
LMArena German12371439
LMArena Japanese11751392
LMArena Korean11731395
LMArena Russian12531432
LMArena Spanish12611456

Instruction Following Hy3 leads

Claude 3.5 Haiku: 62.9 (#234), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Instruction Following12411426
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Hy3 leads

Claude 3.5 Haiku: 38.3 (#200), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Longer Query12611442

Writing & Preference Hy3 leads

Claude 3.5 Haiku: 42.7 (#234), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuHy3
LMArena Text12551439
LMArena Creative Writing12331402
LMArena Multi-Turn12651436
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Hy3?

Hy3 is the stronger model overall, scoring 44.2 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Hy3 share?

17 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper