Model comparison

ERNIE 5.1 vs Gemma 4 31B IT

ERNIE 5.1 and Gemma 4 31B IT score almost the same on the Noometry Index (43.8 vs 43.5), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Summary

  • They share 15 benchmarks with published results for both. ERNIE 5.1 scores higher in 6 categories and Gemma 4 31B IT in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 4 31B IT leads 27.2 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 70.6% for Gemma 4 31B IT.
  • Gemma 4 31B IT has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Gemma 4 31B IT specifications
ERNIE 5.1Gemma 4 31B IT
ProviderBaiduGoogle
Noometry Index43.843.5
Released—2026-04-02
WeightsProprietaryOpen
Context window—262K
Max output—33K
Input $ / M tokens—$0.09
Output $ / M tokens—$0.34
Results tracked1935

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Gemma 4 31B IT: 42.3 (#108)

Coding benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Coding14881459
LMArena WebDev—1366
SciCode—43.4%
WeirdML—52.3%
ALE-Bench—925.5

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Gemma 4 31B IT: —

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Search1227—

Reasoning Gemma 4 31B IT leads

ERNIE 5.1: 21.9 (#211), Gemma 4 31B IT: 27.2 (#122)

Reasoning benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
NYT Connections (extended)23.4%70.6%
LMArena Hard Prompts14811448
Kagi LLM Benchmark—63.5%
CritPt—1.4%
Chess Puzzles—5%
Thematic Generalization—53%
DTBench—82.7%
LMCA—39.3%
Surface Evolver Bench—30.6%
Epoch Capabilities Index—142.74

Math Gemma 4 31B IT leads

ERNIE 5.1: 40.3 (#92), Gemma 4 31B IT: 43.2 (#81)

Math benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Math14811465
OTIS Mock AIME 2024-2025—73.3%

Knowledge ERNIE 5.1 leads

ERNIE 5.1: 41.9 (#102), Gemma 4 31B IT: 37.9 (#151)

Knowledge benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Expert14921465
GPQA Diamond—75.8%
SimpleQA Verified—10.4%
Vectara Hallucination Rate—7.4%

Multimodal Not comparable

ERNIE 5.1: —, Gemma 4 31B IT: 41.6 (#34)

Multimodal benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Vision—1277
LMArena Document—1425

Multilingual ERNIE 5.1 leads

ERNIE 5.1: 55.5 (#29), Gemma 4 31B IT: 53.8 (#57)

Multilingual benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Non-English14541431
LMArena Chinese15081476
LMArena French14881435
LMArena Russian14591460
LMArena Spanish14731444
LMArena German1470—
LMArena Japanese1422—
LMArena Korean1427—

Instruction Following ERNIE 5.1 leads

ERNIE 5.1: 76.7 (#37), Gemma 4 31B IT: 75.5 (#61)

Instruction Following benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Instruction Following14601433

Long Context Too close to call

ERNIE 5.1: 44.7 (#59), Gemma 4 31B IT: 44.2 (#71)

Long Context benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Longer Query14621446

Writing & Preference ERNIE 5.1 leads

ERNIE 5.1: 65.1 (#52), Gemma 4 31B IT: 60.5 (#96)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Gemma 4 31B IT
LMArena Text14681443
LMArena Creative Writing14411415
LMArena Multi-Turn14711452
EQ-Bench Creative Writing—1368
EQ-Bench 4—1120

Frequently asked questions

Is ERNIE 5.1 better than Gemma 4 31B IT?

ERNIE 5.1 and Gemma 4 31B IT score almost the same on the Noometry Index (43.8 vs 43.5), so choose on price, context window or the category you care about most.

Is ERNIE 5.1 or Gemma 4 31B IT better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 42.3 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Gemma 4 31B IT share?

15 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Gemma 4 31B IT has 35.

Related comparisons

Go deeper