Model comparison

DeepSeek-V3 vs Hunyuan Large 2025 02 10

DeepSeek-V3 and Hunyuan Large 2025 02 10 score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Hunyuan Large 2025 02 10 Tencent

38.6

Rank #184 Confirmed

Summary

  • They share 12 benchmarks with published results for both. DeepSeek-V3 scores higher in 5 categories and Hunyuan Large 2025 02 10 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3 leads 57.4 to 48.7.
  • DeepSeek-V3 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3 and Hunyuan Large 2025 02 10 specifications
DeepSeek-V3Hunyuan Large 2025 02 10
ProviderDeepSeekTencent
Noometry Index39.538.6
Released2024-12-26—
WeightsOpenProprietary
Context window164K—
Max output164K—
Input $ / M tokens$0.24—
Output $ / M tokens$0.90—
Results tracked6012

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Hunyuan Large 2025 02 10: 38.2 (#181)

Coding benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Coding13681307
Aider Polyglot55.1%—
SciCode35.8%—
WeirdML36.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Hunyuan Large 2025 02 10: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
METR Time Horizons49.6%—

Reasoning Hunyuan Large 2025 02 10 leads

DeepSeek-V3: 20.5 (#236), Hunyuan Large 2025 02 10: 25.5 (#148)

Reasoning benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Hard Prompts13651286
SimpleBench27.2%—
Kagi LLM Benchmark52.3%—
CritPt0%—
LiveBench Reasoning65.8%—
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
BIG-Bench Hard87.5%—
Epoch Capabilities Index135.94—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Hunyuan Large 2025 02 10 leads

DeepSeek-V3: 32.1 (#219), Hunyuan Large 2025 02 10: 35.8 (#178)

Math benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Math13731281
OTIS Mock AIME 2024-202537.8%—
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge DeepSeek-V3 leads

DeepSeek-V3: 37.5 (#155), Hunyuan Large 2025 02 10: 35.1 (#188)

Knowledge benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Expert13511276
GPQA Diamond67.6%—
MMLU-Pro72.3%—
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), Hunyuan Large 2025 02 10: 42.0 (#200)

Multilingual benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Non-English13581265
LMArena Chinese13911346
LMArena Russian13731266
LMArena French1385—
LMArena German1374—
LMArena Japanese1333—
LMArena Korean1319—
LMArena Spanish1358—

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), Hunyuan Large 2025 02 10: 67.3 (#197)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Instruction Following13451277
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context Hunyuan Large 2025 02 10 leads

DeepSeek-V3: 34.0 (#253), Hunyuan Large 2025 02 10: 40.8 (#149)

Long Context benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Longer Query13521341
Fiction.LiveBench50%—

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), Hunyuan Large 2025 02 10: 48.7 (#197)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Hunyuan Large 2025 02 10
LMArena Text13751288
LMArena Creative Writing13641264
LMArena Multi-Turn13891284
Short-Story Creative Writing77%—
EQ-Bench Creative Writing1472—
WildBench83%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Hunyuan Large 2025 02 10?

DeepSeek-V3 and Hunyuan Large 2025 02 10 score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Is DeepSeek-V3 or Hunyuan Large 2025 02 10 better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 38.2 in the Noometry coding category.

How many benchmarks do DeepSeek-V3 and Hunyuan Large 2025 02 10 share?

12 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Hunyuan Large 2025 02 10 has 12.

Related comparisons

Go deeper