Model comparison

Claude Opus 4.1 vs Hy4 preview

Hy4 preview is the stronger model overall, scoring 45.3 to 41.0 on the Noometry Index.

Last verified . 1 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

Hy4 preview Tencent

45.3

Rank #73 Reported

Summary

  • They share 1 benchmark with published results for both. Claude Opus 4.1 scores higher in 1 category and Hy4 preview in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy4 preview leads 55.7 to 22.3.
  • Hy4 preview is cheaper at $0.83 / $2.50 per million input/output tokens, against $15 / $75 for Claude Opus 4.1.
  • Hy4 preview accepts more context: 1.05M tokens versus 200K.
  • Hy4 preview has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.1 and Hy4 preview specifications
Claude Opus 4.1Hy4 preview
ProviderAnthropicTencent
Noometry Index41.045.3
Released2025-08-052026-08-28
WeightsProprietaryOpen
Context window200K1.05M
Max output32K64K
Input $ / M tokens$15$0.83
Output $ / M tokens$75$2.50
Results tracked483

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Claude Opus 4.1: 44.4 (#73), Hy4 preview: 51.6 (#38)

Coding benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
LMArena WebDev13901632
SWE-bench Verified73.3%—
WeirdML45.9%—
LMArena Coding1479—
ALE-Bench674.77—
AlgoTune1.34—

Agentic & Tool Use Not comparable

Claude Opus 4.1: 35.0 (#41), Hy4 preview: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
Terminal-Bench38%—
GDPval43.6%—
Cybench42%—
DeepResearch Bench48.3%—
LMArena Search1148—
METR Time Horizons66.8%—

Reasoning Too close to call

Claude Opus 4.1: 32.2 (#76), Hy4 preview: 31.9 (#79)

Reasoning benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
SimpleBench60%—
NYT Connections (extended)—68.2%
Chess Puzzles7%—
EnigmaEval7.2%—
EBR-Bench7.9%—
LMArena Hard Prompts1443—
Mystery Game Puzzles21%—
DTBench80%—
LMCA37.1%—
Epoch Capabilities Index144.12—
ForecastBench62—

Math Hy4 preview leads

Claude Opus 4.1: 22.3 (#277), Hy4 preview: 55.7 (#42)

Math benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
FrontierMath (Tiers 1-3)12.6%—
FrontierMath Tier 42.4%—
OTIS Mock AIME 2024-202568.9%—
ProofBench—75%
LMArena Math1431—
FrontierMath (Feb 2025 set)7.2%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Not comparable

Claude Opus 4.1: 42.0 (#101), Hy4 preview: —

Knowledge benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
GPQA Diamond77.3%—
Humanity's Last Exam11.5%—
Confabulations17.1%—
Vectara Hallucination Rate11.8%—
LMArena Expert1439—

Multimodal Not comparable

Claude Opus 4.1: 26.8 (#119), Hy4 preview: —

Multimodal benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
VPCT35%—

Multilingual Not comparable

Claude Opus 4.1: 52.0 (#95), Hy4 preview: —

Multilingual benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
LMArena Non-English1405—
LMArena Chinese1427—
LMArena French1431—
LMArena German1413—
LMArena Japanese1378—
LMArena Korean1380—
LMArena Russian1422—
LMArena Spanish1448—

Instruction Following Not comparable

Claude Opus 4.1: 75.6 (#58), Hy4 preview: —

Instruction Following benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
LMArena Instruction Following1435—

Long Context Not comparable

Claude Opus 4.1: 44.5 (#63), Hy4 preview: —

Long Context benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
LMArena Longer Query1455—

Writing & Preference Not comparable

Claude Opus 4.1: 62.4 (#74), Hy4 preview: —

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1Hy4 preview
LMArena Text1419—
LMArena Creative Writing1412—
Short-Story Creative Writing84.7%—
LMArena Multi-Turn1444—

Frequently asked questions

Is Claude Opus 4.1 better than Hy4 preview?

Hy4 preview is the stronger model overall, scoring 45.3 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4.1 or Hy4 preview?

Hy4 preview is cheaper. It lists at $0.83 per million input tokens and $2.50 per million output tokens; Claude Opus 4.1 lists at $15 and $75.

Is Claude Opus 4.1 or Hy4 preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 44.4 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 200K.

How many benchmarks do Claude Opus 4.1 and Hy4 preview share?

1 benchmark has published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and Hy4 preview has 3.

Related comparisons

Go deeper