Model comparison

ERNIE 5.1 vs Gemini 2.5 Pro

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 43.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Gemini 2.5 Pro Google

45.0

Rank #75 Confirmed

Summary

  • They share 18 benchmarks with published results for both. ERNIE 5.1 scores higher in 5 categories and Gemini 2.5 Pro in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Gemini 2.5 Pro leads 59.8 to 44.7.

Side by side

ERNIE 5.1 and Gemini 2.5 Pro specifications
ERNIE 5.1Gemini 2.5 Pro
ProviderBaiduGoogle
Noometry Index43.845.0
Released—2025-03-25
WeightsProprietaryProprietary
Context window—1.05M
Max output—66K
Input $ / M tokens—$1.25
Output $ / M tokens—$10
Results tracked1978

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Gemini 2.5 Pro: 42.4 (#101)

Coding benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Coding14881452
SWE-bench Verified—57.6%
SWE-bench Verified (bash only)—53.6%
Aider Polyglot—83.1%
LMArena WebDev—1227
SciCode—42.8%
GSO—3.9%
WeirdML—54%
LiveBench Coding—85.9%
CadEval—64%
ALE-Bench—785.52
AlgoTune—1.51

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Gemini 2.5 Pro: 29.2 (#88)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Search12271142
Terminal-Bench—32.6%
GDPval—23.3%
Remote Labor Index—0.8%
TheAgentCompany—30.3%
τ²-bench Banking—13.7%
DeepResearch Bench—42.8%
BALROG—43.3%
METR Time Horizons—55.4%
Vending-Bench 2—573.64

Reasoning Gemini 2.5 Pro leads

ERNIE 5.1: 21.9 (#211), Gemini 2.5 Pro: 28.8 (#99)

Reasoning benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Hard Prompts14811455
ARC-AGI-2—4.9%
SimpleBench—62.4%
Kagi LLM Benchmark—70.3%
NYT Connections (extended)23.4%—
ARC-AGI-1—41%
CritPt—2%
Chess Puzzles—20%
EnigmaEval—5.6%
LiveBench Reasoning—89.8%
DTBench—82.4%
LiveBench Data Analysis—79.9%
LMCA—34.8%
Epoch Capabilities Index—145.32
ForecastBench—61.3
LiveBench—82.3%

Math ERNIE 5.1 leads

ERNIE 5.1: 40.3 (#92), Gemini 2.5 Pro: 32.5 (#213)

Math benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Math14811450
FrontierMath (Tiers 1-3)—24.6%
FrontierMath Tier 4—0%
OTIS Mock AIME 2024-2025—84.7%
Omni-MATH—41.6%
LiveBench Math—90.2%
MATH Level 5—95.9%
FrontierMath (Feb 2025 set)—14.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Gemini 2.5 Pro leads

ERNIE 5.1: 41.9 (#102), Gemini 2.5 Pro: 56.0 (#46)

Knowledge benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Expert14921452
GPQA Diamond—85.3%
Humanity's Last Exam—21.6%
MMLU-Pro—86.3%
Confabulations—10.6%
Vectara Hallucination Rate—7%
GPQA (HELM)—74.9%

Multimodal Not comparable

ERNIE 5.1: —, Gemini 2.5 Pro: 45.2 (#18)

Multimodal benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Vision—1263
GeoBench—86%
VPCT—48%
LMArena Document—1421
SpatialViz-Bench—44.7%

Multilingual Too close to call

ERNIE 5.1: 55.5 (#29), Gemini 2.5 Pro: 55.3 (#31)

Multilingual benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Non-English14541451
LMArena Chinese15081507
LMArena French14881472
LMArena German14701487
LMArena Japanese14221461
LMArena Korean14271434
LMArena Russian14591461
LMArena Spanish14731473

Instruction Following ERNIE 5.1 leads

ERNIE 5.1: 76.7 (#37), Gemini 2.5 Pro: 75.0 (#75)

Instruction Following benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Instruction Following14601437
LiveBench Instruction Following—80.6%
IFEval—84%

Long Context Gemini 2.5 Pro leads

ERNIE 5.1: 44.7 (#59), Gemini 2.5 Pro: 59.8 (#5)

Long Context benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Longer Query14621449
Fiction.LiveBench—91.7%

Writing & Preference ERNIE 5.1 leads

ERNIE 5.1: 65.1 (#52), Gemini 2.5 Pro: 63.7 (#62)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Gemini 2.5 Pro
LMArena Text14681458
LMArena Creative Writing14411454
LMArena Multi-Turn14711453
Short-Story Creative Writing—83.8%
EQ-Bench Creative Writing—1421
WildBench—85.7%
LiveBench Language—67.8%

Frequently asked questions

Is ERNIE 5.1 better than Gemini 2.5 Pro?

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 43.8 on the Noometry Index.

Is ERNIE 5.1 or Gemini 2.5 Pro better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 42.4 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Gemini 2.5 Pro share?

18 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Gemini 2.5 Pro has 78.

Related comparisons

Go deeper