Model comparison

DeepSeek LLM 67B vs Hy4 preview

Hy4 preview is the stronger model overall, scoring 45.3 to 24.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Hy4 preview Tencent

45.3

Rank #73 Reported

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 8.7.

Side by side

DeepSeek LLM 67B and Hy4 preview specifications
DeepSeek LLM 67BHy4 preview
ProviderDeepSeekTencent
Noometry Index24.945.3
Released2023-11-292026-08-28
WeightsOpenOpen
Context window—1.05M
Max output—64K
Input $ / M tokens—$0.75
Output $ / M tokens—$2.25
Results tracked153

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

DeepSeek LLM 67B: 31.9 (#278), Hy4 preview: 51.6 (#38)

Coding benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
LMArena WebDev—1632
LMArena Coding1096—

Reasoning Hy4 preview leads

DeepSeek LLM 67B: 16.5 (#304), Hy4 preview: 31.9 (#79)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
NYT Connections (extended)—68.2%
Chess Puzzles0%—
LMArena Hard Prompts1070—
Epoch Capabilities Index110.5—

Math Hy4 preview leads

DeepSeek LLM 67B: 8.7 (#324), Hy4 preview: 55.7 (#42)

Math benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
OTIS Mock AIME 2024-20250.8%—
ProofBench—75%
LMArena Math1108—
MATH Level 56.4%—

Knowledge Not comparable

DeepSeek LLM 67B: 7.0 (#313), Hy4 preview: —

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
GPQA Diamond24.6%—

Multilingual Not comparable

DeepSeek LLM 67B: 29.4 (#267), Hy4 preview: —

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
LMArena Non-English1073—
LMArena Chinese1132—

Instruction Following Not comparable

DeepSeek LLM 67B: 55.4 (#277), Hy4 preview: —

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
LMArena Instruction Following1079—

Long Context Not comparable

DeepSeek LLM 67B: 33.1 (#265), Hy4 preview: —

Long Context benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
LMArena Longer Query1092—

Writing & Preference Not comparable

DeepSeek LLM 67B: 31.6 (#282), Hy4 preview: —

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BHy4 preview
LMArena Text1105—
LMArena Creative Writing1067—
LMArena Multi-Turn1082—

Frequently asked questions

Is DeepSeek LLM 67B better than Hy4 preview?

Hy4 preview is the stronger model overall, scoring 45.3 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or Hy4 preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Hy4 preview share?

0 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Hy4 preview has 3.

Related comparisons

Go deeper