Model comparison

Claude 3.5 Sonnet vs Hunyuan Large 2025 02 10

Hunyuan Large 2025 02 10 is the stronger model overall, scoring 38.6 to 34.6 on the Noometry Index.

Last verified . 12 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Hunyuan Large 2025 02 10 Tencent

38.6

Rank #184 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 4 categories and Hunyuan Large 2025 02 10 in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hunyuan Large 2025 02 10 leads 35.8 to 19.2.

Side by side

Claude 3.5 Sonnet and Hunyuan Large 2025 02 10 specifications
Claude 3.5 SonnetHunyuan Large 2025 02 10
ProviderAnthropicTencent
Noometry Index34.638.6
Released2024-06-20—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6012

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), Hunyuan Large 2025 02 10: 38.2 (#181)

Coding benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Coding13421307
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Hunyuan Large 2025 02 10: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Hunyuan Large 2025 02 10 leads

Claude 3.5 Sonnet: 23.1 (#183), Hunyuan Large 2025 02 10: 25.5 (#148)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Hard Prompts13051286
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Hunyuan Large 2025 02 10 leads

Claude 3.5 Sonnet: 19.2 (#288), Hunyuan Large 2025 02 10: 35.8 (#178)

Math benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Math13071281
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Hunyuan Large 2025 02 10 leads

Claude 3.5 Sonnet: 28.6 (#245), Hunyuan Large 2025 02 10: 35.1 (#188)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Expert12651276
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Hunyuan Large 2025 02 10: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 43.2 (#185), Hunyuan Large 2025 02 10: 42.0 (#200)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Non-English12831265
LMArena Chinese12721346
LMArena Russian13061266
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Spanish1290—

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Hunyuan Large 2025 02 10: 67.3 (#197)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Instruction Following12971277
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Hunyuan Large 2025 02 10: 40.8 (#149)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Longer Query13111341

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Hunyuan Large 2025 02 10: 48.7 (#197)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetHunyuan Large 2025 02 10
LMArena Text12981288
LMArena Creative Writing12921264
LMArena Multi-Turn13261284
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Hunyuan Large 2025 02 10?

Hunyuan Large 2025 02 10 is the stronger model overall, scoring 38.6 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Hunyuan Large 2025 02 10 better for coding?

They score almost the same on coding (39.0 vs 38.2); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and Hunyuan Large 2025 02 10 share?

12 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Hunyuan Large 2025 02 10 has 12.

Related comparisons

Go deeper