Model comparison

ERNIE 5.1 vs Llama 3.2 90B

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • The widest gap is in math, where ERNIE 5.1 leads 40.3 to 11.1.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Llama 3.2 90B specifications
ERNIE 5.1Llama 3.2 90B
ProviderBaiduMeta
Noometry Index43.827.5
Released—2024-09-24
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked199

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

ERNIE 5.1: 44.0 (#76), Llama 3.2 90B: —

Coding benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
LMArena Coding1488—

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
BALROG—27.3%
LMArena Search1227—

Reasoning Too close to call

ERNIE 5.1: 21.9 (#211), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
NYT Connections (extended)23.4%—
EnigmaEval—0.4%
LMArena Hard Prompts1481—
Epoch Capabilities Index—125.5

Math ERNIE 5.1 leads

ERNIE 5.1: 40.3 (#92), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
OTIS Mock AIME 2024-2025—2.6%
LMArena Math1481—
MATH Level 5—39.4%

Knowledge ERNIE 5.1 leads

ERNIE 5.1: 41.9 (#102), Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
GPQA Diamond—41%
LMArena Expert1492—
MMLU—80.3%

Multimodal Not comparable

ERNIE 5.1: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

ERNIE 5.1: 55.5 (#29), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
LMArena Non-English1454—
LMArena Chinese1508—
LMArena French1488—
LMArena German1470—
LMArena Japanese1422—
LMArena Korean1427—
LMArena Russian1459—
LMArena Spanish1473—

Instruction Following Not comparable

ERNIE 5.1: 76.7 (#37), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
LMArena Instruction Following1460—

Long Context Not comparable

ERNIE 5.1: 44.7 (#59), Llama 3.2 90B: —

Long Context benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
LMArena Longer Query1462—

Writing & Preference Not comparable

ERNIE 5.1: 65.1 (#52), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkERNIE 5.1Llama 3.2 90B
LMArena Text1468—
LMArena Creative Writing1441—
LMArena Multi-Turn1471—

Frequently asked questions

Is ERNIE 5.1 better than Llama 3.2 90B?

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 27.5 on the Noometry Index.

How many benchmarks do ERNIE 5.1 and Llama 3.2 90B share?

0 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper