Model comparison

Gemma 2B vs Hunyuan T1 20250711

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 29.6 on the Noometry Index.

Last verified . 11 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Hunyuan T1 20250711 Tencent

42.5

Rank #114 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Gemma 2B scores higher in 0 categories and Hunyuan T1 20250711 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hunyuan T1 20250711 leads 59.5 to 24.0.
  • Gemma 2B has downloadable open weights; the other is API-only.

Side by side

Gemma 2B and Hunyuan T1 20250711 specifications
Gemma 2BHunyuan T1 20250711
ProviderGoogleTencent
Noometry Index29.642.5
Released2024-02-21—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2313

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hunyuan T1 20250711 leads

Gemma 2B: 29.4 (#305), Hunyuan T1 20250711: 40.9 (#129)

Coding benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Coding10101390
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Hunyuan T1 20250711 leads

Gemma 2B: 18.8 (#275), Hunyuan T1 20250711: 28.5 (#103)

Reasoning benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Hard Prompts9891399
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Hunyuan T1 20250711 leads

Gemma 2B: 30.0 (#239), Hunyuan T1 20250711: 38.7 (#130)

Math benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Math10091414
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Hunyuan T1 20250711: 38.8 (#141)

Knowledge benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Expert—1395
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multilingual Hunyuan T1 20250711 leads

Gemma 2B: 23.0 (#294), Hunyuan T1 20250711: 51.2 (#112)

Multilingual benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Non-English9581395
LMArena Chinese9861425
LMArena Russian9371385
LMArena Korean—1406

Instruction Following Hunyuan T1 20250711 leads

Gemma 2B: 48.5 (#302), Hunyuan T1 20250711: 72.6 (#138)

Instruction Following benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Instruction Following9701374

Long Context Hunyuan T1 20250711 leads

Gemma 2B: 29.9 (#291), Hunyuan T1 20250711: 42.2 (#128)

Long Context benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Longer Query9811384

Writing & Preference Hunyuan T1 20250711 leads

Gemma 2B: 24.0 (#308), Hunyuan T1 20250711: 59.5 (#109)

Writing & Preference benchmarks
BenchmarkGemma 2BHunyuan T1 20250711
LMArena Text10021401
LMArena Creative Writing9871392
LMArena Multi-Turn9451393

Frequently asked questions

Is Gemma 2B better than Hunyuan T1 20250711?

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 29.6 on the Noometry Index.

Is Gemma 2B or Hunyuan T1 20250711 better for coding?

Hunyuan T1 20250711 scores higher on coding benchmarks: 40.9 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and Hunyuan T1 20250711 share?

11 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Hunyuan T1 20250711 has 13.

Related comparisons

Go deeper