Model comparison

Claude 3.5 Haiku vs Hy4 preview

Hy4 preview is the stronger model overall, scoring 45.3 to 29.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Hy4 preview Tencent

45.3

Rank #73 Reported

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 14.7.
  • Hy4 preview has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Hy4 preview specifications
Claude 3.5 HaikuHy4 preview
ProviderAnthropicTencent
Noometry Index29.245.3
Released2024-10-222026-08-28
WeightsProprietaryOpen
Context window—1.05M
Max output—64K
Input $ / M tokens—$0.75
Output $ / M tokens—$2.25
Results tracked493

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Claude 3.5 Haiku: 32.9 (#265), Hy4 preview: 51.6 (#38)

Coding benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
Aider Polyglot28%—
LMArena WebDev—1632
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
LMArena Coding1286—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Hy4 preview: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
BALROG19.3%—

Reasoning Hy4 preview leads

Claude 3.5 Haiku: 17.7 (#290), Hy4 preview: 31.9 (#79)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
NYT Connections (extended)—68.2%
CritPt0%—
LiveBench Reasoning28.1%—
LMArena Hard Prompts1251—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Hy4 preview leads

Claude 3.5 Haiku: 14.7 (#300), Hy4 preview: 55.7 (#42)

Math benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
OTIS Mock AIME 2024-20254.3%—
ProofBench—75%
Omni-MATH22.4%—
LiveBench Math35.5%—
LMArena Math1244—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Claude 3.5 Haiku: 18.7 (#281), Hy4 preview: —

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
LMArena Expert1208—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Hy4 preview: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
LMArena Vision1092—
GeoBench34%—

Multilingual Not comparable

Claude 3.5 Haiku: 40.0 (#218), Hy4 preview: —

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
LMArena Non-English1238—
LMArena Chinese1229—
LMArena French1264—
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Russian1253—
LMArena Spanish1261—

Instruction Following Not comparable

Claude 3.5 Haiku: 62.9 (#234), Hy4 preview: —

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
LiveBench Instruction Following61.9%—
IFEval79.2%—
LMArena Instruction Following1241—

Long Context Not comparable

Claude 3.5 Haiku: 38.3 (#200), Hy4 preview: —

Long Context benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
LMArena Longer Query1261—

Writing & Preference Not comparable

Claude 3.5 Haiku: 42.7 (#234), Hy4 preview: —

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuHy4 preview
LMArena Text1255—
LMArena Creative Writing1233—
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LMArena Multi-Turn1265—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Hy4 preview?

Hy4 preview is the stronger model overall, scoring 45.3 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Hy4 preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Hy4 preview share?

0 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Hy4 preview has 3.

Related comparisons

Go deeper