Model comparison

Gemma 4 31B IT vs Hy3

Gemma 4 31B IT and Hy3 score almost the same on the Noometry Index (43.5 vs 44.2), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Hy3 Tencent

44.2

Rank #79 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Gemma 4 31B IT scores higher in 5 categories and Hy3 in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Hy3 leads 46.8 to 42.3.
  • The biggest single-benchmark swing is NYT Connections (extended): 70.6% for Gemma 4 31B IT and 41.2% for Hy3.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $0.13 / $0.53 for Hy3.

Side by side

Gemma 4 31B IT and Hy3 specifications
Gemma 4 31B ITHy3
ProviderGoogleTencent
Noometry Index43.544.2
Released2026-04-022026-07-06
WeightsOpenOpen
Context window262K262K
Max output33K128K
Input $ / M tokens$0.09$0.13
Output $ / M tokens$0.34$0.53
Results tracked3519

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Gemma 4 31B IT: 42.3 (#108), Hy3: 46.8 (#63)

Coding benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena WebDev13661508
LMArena Coding14591464
SciCode43.4%—
WeirdML52.3%—
ALE-Bench925.5—

Reasoning Gemma 4 31B IT leads

Gemma 4 31B IT: 27.2 (#122), Hy3: 26.1 (#136)

Reasoning benchmarks
BenchmarkGemma 4 31B ITHy3
NYT Connections (extended)70.6%41.2%
LMArena Hard Prompts14481447
Kagi LLM Benchmark63.5%—
CritPt1.4%—
Chess Puzzles5%—
Thematic Generalization53%—
DTBench82.7%—
LMCA39.3%—
Surface Evolver Bench30.6%—
Epoch Capabilities Index142.74—

Math Gemma 4 31B IT leads

Gemma 4 31B IT: 43.2 (#81), Hy3: 40.1 (#93)

Math benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Math14651475
OTIS Mock AIME 2024-202573.3%—

Knowledge Hy3 leads

Gemma 4 31B IT: 37.9 (#151), Hy3: 40.8 (#114)

Knowledge benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Expert14651460
GPQA Diamond75.8%—
SimpleQA Verified10.4%—
Vectara Hallucination Rate7.4%—

Multimodal Not comparable

Gemma 4 31B IT: 41.6 (#34), Hy3: —

Multimodal benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Vision1277—
LMArena Document1425—

Multilingual Too close to call

Gemma 4 31B IT: 53.8 (#57), Hy3: 53.5 (#65)

Multilingual benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Non-English14311426
LMArena Chinese14761493
LMArena French14351461
LMArena Russian14601432
LMArena Spanish14441456
LMArena German—1439
LMArena Japanese—1392
LMArena Korean—1395

Instruction Following Too close to call

Gemma 4 31B IT: 75.5 (#61), Hy3: 75.1 (#70)

Instruction Following benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Instruction Following14331426

Long Context Too close to call

Gemma 4 31B IT: 44.2 (#71), Hy3: 44.1 (#75)

Long Context benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Longer Query14461442

Writing & Preference Hy3 leads

Gemma 4 31B IT: 60.5 (#96), Hy3: 62.2 (#81)

Writing & Preference benchmarks
BenchmarkGemma 4 31B ITHy3
LMArena Text14431439
LMArena Creative Writing14151402
LMArena Multi-Turn14521436
EQ-Bench Creative Writing1368—
EQ-Bench 41120—

Frequently asked questions

Is Gemma 4 31B IT better than Hy3?

Gemma 4 31B IT and Hy3 score almost the same on the Noometry Index (43.5 vs 44.2), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 4 31B IT or Hy3?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; Hy3 lists at $0.13 and $0.53.

Is Gemma 4 31B IT or Hy3 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 42.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Gemma 4 31B IT and Hy3 share?

16 benchmarks have published results for both models. Gemma 4 31B IT has 35 scored results on Noometry and Hy3 has 19.

Related comparisons

Go deeper