Model comparison

Hy4 preview vs Qwen3.5 122B-A10B

Hy4 preview is the stronger model overall, scoring 45.3 to 42.1 on the Noometry Index.

Last verified . 2 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Qwen3.5 122B-A10B Alibaba (Qwen)

42.1

Rank #119 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Hy4 preview scores higher in 3 categories and Qwen3.5 122B-A10B in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy4 preview leads 55.7 to 39.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 68.2% for Hy4 preview and 51.7% for Qwen3.5 122B-A10B.
  • Both cost about the same: $0.75 input and $2.25 output per million tokens.
  • Hy4 preview accepts more context: 1.05M tokens versus 262K.

Side by side

Hy4 preview and Qwen3.5 122B-A10B specifications
Hy4 previewQwen3.5 122B-A10B
ProviderTencentAlibaba (Qwen)
Noometry Index45.342.1
Released2026-08-282026-02-23
WeightsOpenOpen
Context window1.05M262K
Max output64K66K
Input $ / M tokens$0.75$0.40
Output $ / M tokens$2.25$3.20
Results tracked327

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Qwen3.5 122B-A10B: 39.1 (#162)

Coding benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
LMArena WebDev16321360
SciCode—35.6%
LMArena Coding—1436

Reasoning Hy4 preview leads

Hy4 preview: 31.9 (#79), Qwen3.5 122B-A10B: 27.2 (#123)

Reasoning benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
NYT Connections (extended)68.2%51.7%
CritPt—0.9%
Thematic Generalization—51.2%
LMArena Hard Prompts—1421
Mystery Game Puzzles—17%
DTBench—84.3%
LMCA—32.2%

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Qwen3.5 122B-A10B: 39.1 (#112)

Math benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
ProofBench75%—
LMArena Math—1432

Knowledge Not comparable

Hy4 preview: —, Qwen3.5 122B-A10B: 38.8 (#142)

Knowledge benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
Vectara Hallucination Rate—11.2%
LMArena Expert—1432

Multimodal Not comparable

Hy4 preview: —, Qwen3.5 122B-A10B: 39.6 (#57)

Multimodal benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
LMArena Vision—1245

Multilingual Not comparable

Hy4 preview: —, Qwen3.5 122B-A10B: 51.6 (#107)

Multilingual benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
LMArena Non-English—1400
LMArena Chinese—1462
LMArena French—1442
LMArena German—1426
LMArena Japanese—1367
LMArena Korean—1352
LMArena Russian—1400
LMArena Spanish—1424

Instruction Following Not comparable

Hy4 preview: —, Qwen3.5 122B-A10B: 73.8 (#115)

Instruction Following benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
LMArena Instruction Following—1399

Long Context Not comparable

Hy4 preview: —, Qwen3.5 122B-A10B: 43.0 (#109)

Long Context benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
LMArena Longer Query—1410

Writing & Preference Not comparable

Hy4 preview: —, Qwen3.5 122B-A10B: 60.0 (#105)

Writing & Preference benchmarks
BenchmarkHy4 previewQwen3.5 122B-A10B
LMArena Text—1417
LMArena Creative Writing—1368
LMArena Multi-Turn—1416

Frequently asked questions

Is Hy4 preview better than Qwen3.5 122B-A10B?

Hy4 preview is the stronger model overall, scoring 45.3 to 42.1 on the Noometry Index.

Which is cheaper, Hy4 preview or Qwen3.5 122B-A10B?

Qwen3.5 122B-A10B is cheaper. It lists at $0.40 per million input tokens and $3.20 per million output tokens; Hy4 preview lists at $0.75 and $2.25.

Is Hy4 preview or Qwen3.5 122B-A10B better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 39.1 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 262K.

How many benchmarks do Hy4 preview and Qwen3.5 122B-A10B share?

2 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Qwen3.5 122B-A10B has 27.

Related comparisons

Go deeper