Model comparison

Hy4 preview vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 45.3 on the Noometry Index. Hy4 preview costs 2.6× less per token, which makes it the better buy when Qwen3.6 Max Preview's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Hy4 preview scores higher in 2 categories and Qwen3.6 Max Preview in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Max Preview leads 41.7 to 31.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 68.2% for Hy4 preview and 74.1% for Qwen3.6 Max Preview.
  • Hy4 preview is cheaper at $0.75 / $2.25 per million input/output tokens, against $1.30 / $7.80 for Qwen3.6 Max Preview.
  • Hy4 preview accepts more context: 1.05M tokens versus 262K.
  • Hy4 preview has downloadable open weights; the other is API-only.

Side by side

Hy4 preview and Qwen3.6 Max Preview specifications
Hy4 previewQwen3.6 Max Preview
ProviderTencentAlibaba (Qwen)
Noometry Index45.351.5
Released2026-08-282026-04-20
WeightsOpenProprietary
Context window1.05M262K
Max output64K66K
Input $ / M tokens$0.75$1.30
Output $ / M tokens$2.25$7.80
Results tracked329

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
LMArena WebDev16321482
SWE-bench Verified—76.7%
LMArena Coding—1471

Agentic & Tool Use Not comparable

Hy4 preview: —, Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Hy4 preview: 31.9 (#79), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
NYT Connections (extended)68.2%74.1%
SimpleBench—63%
Chess Puzzles—20%
LMArena Hard Prompts—1457
Mystery Game Puzzles—19%
DTBench—87.2%
LMCA—42.5%
Epoch Capabilities Index—149.24

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
OTIS Mock AIME 2024-2025—91.1%
ProofBench75%—
LMArena Math—1465
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Not comparable

Hy4 preview: —, Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
GPQA Diamond—87.4%
SimpleQA Verified—52%
LMArena Expert—1478

Multilingual Not comparable

Hy4 preview: —, Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
LMArena Non-English—1437
LMArena Chinese—1487
LMArena French—1449
LMArena Russian—1445
LMArena Spanish—1454

Instruction Following Not comparable

Hy4 preview: —, Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
LMArena Instruction Following—1438

Long Context Not comparable

Hy4 preview: —, Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
LMArena Longer Query—1457

Writing & Preference Not comparable

Hy4 preview: —, Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkHy4 previewQwen3.6 Max Preview
LMArena Text—1447
LMArena Creative Writing—1435
LMArena Multi-Turn—1456

Frequently asked questions

Is Hy4 preview better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 45.3 on the Noometry Index. Hy4 preview costs 2.6× less per token, which makes it the better buy when Qwen3.6 Max Preview's lead doesn't matter for your workload.

Which is cheaper, Hy4 preview or Qwen3.6 Max Preview?

Hy4 preview is cheaper. It lists at $0.75 per million input tokens and $2.25 per million output tokens; Qwen3.6 Max Preview lists at $1.30 and $7.80.

Is Hy4 preview or Qwen3.6 Max Preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 262K.

How many benchmarks do Hy4 preview and Qwen3.6 Max Preview share?

2 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper