Model comparison

Gemini 2.0 Flash (Feb 2025) vs Hunyuan T1 20250711

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 35.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Hunyuan T1 20250711 Tencent

42.5

Rank #114 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 1 category and Hunyuan T1 20250711 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Hunyuan T1 20250711 leads 28.5 to 15.2.

Side by side

Gemini 2.0 Flash (Feb 2025) and Hunyuan T1 20250711 specifications
Gemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
ProviderGoogleTencent
Noometry Index35.142.5
Released2024-12-06—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked5413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hunyuan T1 20250711 leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Hunyuan T1 20250711: 40.9 (#129)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Coding13501390
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
WeirdML25.8%—
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—

Agentic & Tool Use Not comparable

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Hunyuan T1 20250711: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
TheAgentCompany11.4%—

Reasoning Hunyuan T1 20250711 leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Hunyuan T1 20250711: 28.5 (#103)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Hard Prompts13461399
ARC-AGI-21.3%—
SimpleBench31.1%—
Kagi LLM Benchmark37.8%—
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
DTBench63.2%—
LiveBench Data Analysis69.4%—
Epoch Capabilities Index135.36—
LiveBench66.9%—

Math Too close to call

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Hunyuan T1 20250711: 38.7 (#130)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Math13521414
OTIS Mock AIME 2024-202557.8%—
Omni-MATH45.9%—
LiveBench Math75.8%—
MATH Level 582.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Hunyuan T1 20250711 leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Hunyuan T1 20250711: 38.8 (#141)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Expert13391395
GPQA Diamond64.1%—
Humanity's Last Exam6.6%—
MMLU-Pro73.7%—
Confabulations12.4%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Hunyuan T1 20250711: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Vision1158—
GeoBench77%—

Multilingual Hunyuan T1 20250711 leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Hunyuan T1 20250711: 51.2 (#112)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Non-English13421395
LMArena Chinese13731425
LMArena Korean13131406
LMArena Russian13511385
LMArena French1391—
LMArena German1353—
LMArena Japanese1294—
LMArena Spanish1363—

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Hunyuan T1 20250711: 72.6 (#138)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Instruction Following13361374
LiveBench Instruction Following85.8%—
IFEval84.1%—

Long Context Hunyuan T1 20250711 leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Hunyuan T1 20250711: 42.2 (#128)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Longer Query13441384
Fiction.LiveBench61.1%—

Writing & Preference Hunyuan T1 20250711 leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Hunyuan T1 20250711: 59.5 (#109)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Hunyuan T1 20250711
LMArena Text13541401
LMArena Creative Writing13401392
LMArena Multi-Turn13501393
Short-Story Creative Writing73.8%—
EQ-Bench Creative Writing1128—
WildBench80%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Hunyuan T1 20250711?

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 35.1 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or Hunyuan T1 20250711 better for coding?

Hunyuan T1 20250711 scores higher on coding benchmarks: 40.9 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Hunyuan T1 20250711 share?

13 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Hunyuan T1 20250711 has 13.

Related comparisons

Go deeper