Model comparison

Claude Sonnet 4 vs Hunyuan Turbos 20250226

Claude Sonnet 4 and Hunyuan Turbos 20250226 score almost the same on the Noometry Index (40.8 vs 41.3), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Hunyuan Turbos 20250226 Tencent

41.3

Rank #139 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Claude Sonnet 4 scores higher in 4 categories and Hunyuan Turbos 20250226 in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Hunyuan Turbos 20250226 leads 41.6 to 33.7.

Side by side

Claude Sonnet 4 and Hunyuan Turbos 20250226 specifications
Claude Sonnet 4Hunyuan Turbos 20250226
ProviderAnthropicTencent
Noometry Index40.841.3
Released2025-05-22—
WeightsProprietaryProprietary
Context window200K—
Max output64K—
Input $ / M tokens$3—
Output $ / M tokens$15—
Results tracked5816

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude Sonnet 4: 43.5 (#88), Hunyuan Turbos 20250226: 40.0 (#152)

Coding benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Coding14141361
SWE-bench Verified (bash only)64.9%—
Aider Polyglot61.3%—
SciCode40%—
GSO4.9%—
WeirdML46.1%—
ALE-Bench655.35—

Agentic & Tool Use Not comparable

Claude Sonnet 4: 38.5 (#31), Hunyuan Turbos 20250226: —

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
TheAgentCompany33.1%—
Cybench35%—
DeepResearch Bench46.6%—
OSWorld43.9%—
METR Time Horizons62%—

Reasoning Hunyuan Turbos 20250226 leads

Claude Sonnet 4: 22.9 (#187), Hunyuan Turbos 20250226: 27.8 (#113)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Hard Prompts13721374
ARC-AGI-25.9%—
SimpleBench45.5%—
Kagi LLM Benchmark73%—
ARC-AGI-140%—
CritPt0.3%—
EnigmaEval3.1%—
DTBench77.1%—
LMCA29%—
Epoch Capabilities Index141.69—
ForecastBench60.2—

Math Claude Sonnet 4 leads

Claude Sonnet 4: 43.3 (#80), Hunyuan Turbos 20250226: 37.5 (#154)

Math benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Math13751359
OTIS Mock AIME 2024-202571.1%—
Omni-MATH60.2%—
MATH Level 584.4%—
FrontierMath (Feb 2025 set)4.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Claude Sonnet 4 leads

Claude Sonnet 4: 41.8 (#108), Hunyuan Turbos 20250226: 37.0 (#161)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Expert13721339
GPQA Diamond79.2%—
Humanity's Last Exam7.8%—
MMLU-Pro84.3%—
Confabulations13.2%—
Vectara Hallucination Rate10.3%—
GPQA (HELM)70.6%—

Multimodal Not comparable

Claude Sonnet 4: 26.2 (#121), Hunyuan Turbos 20250226: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Vision1191—
GeoBench37%—
VPCT34%—
MindCube44.8%—

Multilingual Hunyuan Turbos 20250226 leads

Claude Sonnet 4: 46.7 (#156), Hunyuan Turbos 20250226: 48.9 (#136)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Non-English13331363
LMArena Chinese13501417
LMArena French13631391
LMArena German13311355
LMArena Japanese13021342
LMArena Korean12911351
LMArena Russian13551368
LMArena Spanish1357—

Instruction Following Too close to call

Claude Sonnet 4: 71.7 (#145), Hunyuan Turbos 20250226: 71.0 (#158)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Instruction Following13761344
IFEval84%—

Long Context Hunyuan Turbos 20250226 leads

Claude Sonnet 4: 33.7 (#259), Hunyuan Turbos 20250226: 41.6 (#136)

Long Context benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Longer Query13981366
Fiction.LiveBench46.9%—

Writing & Preference Too close to call

Claude Sonnet 4: 57.1 (#132), Hunyuan Turbos 20250226: 57.4 (#128)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Hunyuan Turbos 20250226
LMArena Text13511377
LMArena Creative Writing13451359
LMArena Multi-Turn13761387
Short-Story Creative Writing81.4%—
EQ-Bench Creative Writing1483—
WildBench83.8%—

Frequently asked questions

Is Claude Sonnet 4 better than Hunyuan Turbos 20250226?

Claude Sonnet 4 and Hunyuan Turbos 20250226 score almost the same on the Noometry Index (40.8 vs 41.3), so choose on price, context window or the category you care about most.

Is Claude Sonnet 4 or Hunyuan Turbos 20250226 better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 40.0 in the Noometry coding category.

How many benchmarks do Claude Sonnet 4 and Hunyuan Turbos 20250226 share?

16 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Hunyuan Turbos 20250226 has 16.

Related comparisons

Go deeper