Model comparison

ERNIE 5.1 vs Gemini 3.1 Pro Preview

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 43.8 on the Noometry Index.

Last verified . 19 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Summary

  • They share 19 benchmarks with published results for both. ERNIE 5.1 scores higher in 1 category and Gemini 3.1 Pro Preview in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Pro Preview leads 71.7 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 97.4% for Gemini 3.1 Pro Preview.

Side by side

ERNIE 5.1 and Gemini 3.1 Pro Preview specifications
ERNIE 5.1Gemini 3.1 Pro Preview
ProviderBaiduGoogle
Noometry Index43.856.7
Released—2026-02-19
WeightsProprietaryProprietary
Context window—1.05M
Max output—66K
Input $ / M tokens—$2
Output $ / M tokens—$12
Results tracked1971

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Gemini 3.1 Pro Preview: 42.5 (#99)

Coding benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Coding14881484
SWE-bench Verified—75.6%
DeepSWE—11.7%
LMArena WebDev—1447
SciCode—58.9%
GSO—22.6%
WeirdML—72.1%
MirrorCode—8.9%
ALE-Bench—1,161
AlgoTune—2.02

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Gemini 3.1 Pro Preview: 37.7 (#34)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Search12271211
Terminal-Bench—80.2%
APEX-Agents—35.3%
τ²-bench Banking—26%
DeepResearch Bench—47.8%
PostTrainBench—22%
BALROG—57%
ExploitBench—26.1%
GBAEval—0.8%
GDP.pdf—17%
METR Time Horizons—77%
Vending-Bench 2—3,774

Reasoning Gemini 3.1 Pro Preview leads

ERNIE 5.1: 21.9 (#211), Gemini 3.1 Pro Preview: 71.7 (#12)

Reasoning benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
NYT Connections (extended)23.4%97.4%
LMArena Hard Prompts14811485
ARC-AGI-2—77.1%
SimpleBench—79.6%
ARC-AGI-1—98%
CritPt—17.7%
Chess Puzzles—55%
EnigmaEval—36.8%
Thematic Generalization—79.4%
EBR-Bench—14.3%
Mystery Game Puzzles—34%
DTBench—97.1%
LMCA—53.8%
Epoch Capabilities Index—154.77
ForecastBench—59

Math Gemini 3.1 Pro Preview leads

ERNIE 5.1: 40.3 (#92), Gemini 3.1 Pro Preview: 62.1 (#34)

Math benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Math14811485
FrontierMath (Tiers 1-3)—59.6%
FrontierMath Tier 4—26.8%
MathArena Final-Answer Competitions—86.5%
OTIS Mock AIME 2024-2025—95.6%
ProofBench—26%
FrontierMath (Feb 2025 set)—36.9%
FrontierMath Tier 4 (v1)—16.7%

Knowledge Gemini 3.1 Pro Preview leads

ERNIE 5.1: 41.9 (#102), Gemini 3.1 Pro Preview: 71.8 (#3)

Knowledge benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Expert14921485
GPQA Diamond—94.4%
Humanity's Last Exam—46.4%
SimpleQA Verified—73.5%
Vectara Hallucination Rate—10.4%

Multimodal Not comparable

ERNIE 5.1: —, Gemini 3.1 Pro Preview: 37.9 (#69)

Multimodal benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Vision—1296
Blueprint-Bench 2—26.5%
Furniture Assembly—26.7%
LMArena Document—1444

Multilingual Gemini 3.1 Pro Preview leads

ERNIE 5.1: 55.5 (#29), Gemini 3.1 Pro Preview: 57.0 (#12)

Multilingual benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Non-English14541477
LMArena Chinese15081529
LMArena French14881487
LMArena German14701491
LMArena Japanese14221493
LMArena Korean14271455
LMArena Russian14591498
LMArena Spanish14731479

Instruction Following Too close to call

ERNIE 5.1: 76.7 (#37), Gemini 3.1 Pro Preview: 77.0 (#32)

Instruction Following benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Instruction Following14601466

Long Context Gemini 3.1 Pro Preview leads

ERNIE 5.1: 44.7 (#59), Gemini 3.1 Pro Preview: 47.4 (#18)

Long Context benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Longer Query14621483
CL-bench—20.8%
CL-bench Life—16.9%

Writing & Preference Gemini 3.1 Pro Preview leads

ERNIE 5.1: 65.1 (#52), Gemini 3.1 Pro Preview: 66.1 (#37)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Gemini 3.1 Pro Preview
LMArena Text14681481
LMArena Creative Writing14411482
LMArena Multi-Turn14711488
EQ-Bench Creative Writing—1491
EQ-Bench 4—1142

Frequently asked questions

Is ERNIE 5.1 better than Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 43.8 on the Noometry Index.

Is ERNIE 5.1 or Gemini 3.1 Pro Preview better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 42.5 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Gemini 3.1 Pro Preview share?

19 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Gemini 3.1 Pro Preview has 71.

Related comparisons

Go deeper