Model comparison

Claude 2.1 vs Hy3

Hy3 is the stronger model overall, scoring 44.2 to 25.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • The widest gap is in math, where Hy3 leads 40.1 to 10.2.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Hy3 specifications
Claude 2.1Hy3
ProviderAnthropicTencent
Noometry Index25.244.2
Released2023-11-212026-07-06
WeightsProprietaryOpen
Context window—262K
Max output—128K
Input $ / M tokens—$0.13
Output $ / M tokens—$0.53
Results tracked719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Claude 2.1: 26.2 (#327), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkClaude 2.1Hy3
LMArena WebDev—1508
WeirdML7.1%—
LMArena Coding—1464

Reasoning Hy3 leads

Claude 2.1: 21.4 (#221), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkClaude 2.1Hy3
NYT Connections (extended)—41.2%
LMArena Hard Prompts—1447
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Hy3 leads

Claude 2.1: 10.2 (#315), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkClaude 2.1Hy3
OTIS Mock AIME 2024-20251.9%—
LMArena Math—1475

Knowledge Hy3 leads

Claude 2.1: 15.4 (#292), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkClaude 2.1Hy3
GPQA Diamond33%—
LMArena Expert—1460
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkClaude 2.1Hy3
LMArena Non-English—1426
LMArena Chinese—1493
LMArena French—1461
LMArena German—1439
LMArena Japanese—1392
LMArena Korean—1395
LMArena Russian—1432
LMArena Spanish—1456

Instruction Following Not comparable

Claude 2.1: —, Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkClaude 2.1Hy3
LMArena Instruction Following—1426

Long Context Not comparable

Claude 2.1: —, Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkClaude 2.1Hy3
LMArena Longer Query—1442

Writing & Preference Not comparable

Claude 2.1: —, Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkClaude 2.1Hy3
LMArena Text—1439
LMArena Creative Writing—1402
LMArena Multi-Turn—1436

Frequently asked questions

Is Claude 2.1 better than Hy3?

Hy3 is the stronger model overall, scoring 44.2 to 25.2 on the Noometry Index.

Is Claude 2.1 or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Hy3 share?

0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper