Model comparison

Claude 3.5 Sonnet vs Hunyuan T1 20250711

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 34.6 on the Noometry Index.

Last verified . 13 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Hunyuan T1 20250711 Tencent

42.5

Rank #114 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 0 categories and Hunyuan T1 20250711 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hunyuan T1 20250711 leads 38.7 to 19.2.

Side by side

Claude 3.5 Sonnet and Hunyuan T1 20250711 specifications
Claude 3.5 SonnetHunyuan T1 20250711
ProviderAnthropicTencent
Noometry Index34.642.5
Released2024-06-20—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 39.0 (#165), Hunyuan T1 20250711: 40.9 (#129)

Coding benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Coding13421390
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Hunyuan T1 20250711: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 23.1 (#183), Hunyuan T1 20250711: 28.5 (#103)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Hard Prompts13051399
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 19.2 (#288), Hunyuan T1 20250711: 38.7 (#130)

Math benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Math13071414
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 28.6 (#245), Hunyuan T1 20250711: 38.8 (#141)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Expert12651395
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Hunyuan T1 20250711: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 43.2 (#185), Hunyuan T1 20250711: 51.2 (#112)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Non-English12831395
LMArena Chinese12721425
LMArena Korean12001406
LMArena Russian13061385
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Spanish1290—

Instruction Following Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 68.8 (#182), Hunyuan T1 20250711: 72.6 (#138)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Instruction Following12971374
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 39.9 (#167), Hunyuan T1 20250711: 42.2 (#128)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Longer Query13111384

Writing & Preference Hunyuan T1 20250711 leads

Claude 3.5 Sonnet: 52.9 (#164), Hunyuan T1 20250711: 59.5 (#109)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetHunyuan T1 20250711
LMArena Text12981401
LMArena Creative Writing12921392
LMArena Multi-Turn13261393
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Hunyuan T1 20250711?

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Hunyuan T1 20250711 better for coding?

Hunyuan T1 20250711 scores higher on coding benchmarks: 40.9 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Hunyuan T1 20250711 share?

13 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Hunyuan T1 20250711 has 13.

Related comparisons

Go deeper