Model comparison

ERNIE 5.1 vs Olmo 3.1 32b Think

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 37.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Summary

  • They share 15 benchmarks with published results for both. ERNIE 5.1 scores higher in 7 categories and Olmo 3.1 32b Think in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where ERNIE 5.1 leads 65.1 to 46.2.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Olmo 3.1 32b Think specifications
ERNIE 5.1Olmo 3.1 32b Think
ProviderBaiduAllen Institute for AI (Ai2)
Noometry Index43.837.9
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Coding14881291

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Olmo 3.1 32b Think: —

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Search1227—

Reasoning Olmo 3.1 32b Think leads

ERNIE 5.1: 21.9 (#211), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Hard Prompts14811272
NYT Connections (extended)23.4%—

Math ERNIE 5.1 leads

ERNIE 5.1: 40.3 (#92), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Math14811305

Knowledge ERNIE 5.1 leads

ERNIE 5.1: 41.9 (#102), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Expert14921295

Multilingual ERNIE 5.1 leads

ERNIE 5.1: 55.5 (#29), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Non-English14541209
LMArena Chinese15081242
LMArena French14881260
LMArena German14701262
LMArena Russian14591193
LMArena Spanish14731289
LMArena Japanese1422—
LMArena Korean1427—

Instruction Following ERNIE 5.1 leads

ERNIE 5.1: 76.7 (#37), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Instruction Following14601247

Long Context ERNIE 5.1 leads

ERNIE 5.1: 44.7 (#59), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Longer Query14621272

Writing & Preference ERNIE 5.1 leads

ERNIE 5.1: 65.1 (#52), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Olmo 3.1 32b Think
LMArena Text14681272
LMArena Creative Writing14411226
LMArena Multi-Turn14711252

Frequently asked questions

Is ERNIE 5.1 better than Olmo 3.1 32b Think?

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 37.9 on the Noometry Index.

Is ERNIE 5.1 or Olmo 3.1 32b Think better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 37.7 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper