Model comparison

ERNIE 5.1 vs Qwen3 8B

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 33.7 on the Noometry Index.

Last verified . 0 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • The widest gap is in coding, where ERNIE 5.1 leads 44.0 to 34.0.
  • Qwen3 8B has downloadable open weights; the other is API-only.

Side by side

ERNIE 5.1 and Qwen3 8B specifications
ERNIE 5.1Qwen3 8B
ProviderBaiduAlibaba (Qwen)
Noometry Index43.833.7
Released—2025-04
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked1911

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding ERNIE 5.1 leads

ERNIE 5.1: 44.0 (#76), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkERNIE 5.1Qwen3 8B
SciCode—22.6%
LMArena Coding1488—

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Qwen3 8B
Berkeley Function Calling Leaderboard—42.6%
LMArena Search1227—

Reasoning ERNIE 5.1 leads

ERNIE 5.1: 21.9 (#211), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkERNIE 5.1Qwen3 8B
NYT Connections (extended)23.4%—
CritPt—0%
Chess Puzzles—5%
LMArena Hard Prompts1481—
DTBench—59.7%
LMCA—8.8%
Epoch Capabilities Index—136.17

Math ERNIE 5.1 leads

ERNIE 5.1: 40.3 (#92), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkERNIE 5.1Qwen3 8B
OTIS Mock AIME 2024-2025—56.1%
LMArena Math1481—

Knowledge ERNIE 5.1 leads

ERNIE 5.1: 41.9 (#102), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkERNIE 5.1Qwen3 8B
GPQA Diamond—56.8%
Vectara Hallucination Rate—4.8%
LMArena Expert1492—

Multilingual Not comparable

ERNIE 5.1: 55.5 (#29), Qwen3 8B: —

Multilingual benchmarks
BenchmarkERNIE 5.1Qwen3 8B
LMArena Non-English1454—
LMArena Chinese1508—
LMArena French1488—
LMArena German1470—
LMArena Japanese1422—
LMArena Korean1427—
LMArena Russian1459—
LMArena Spanish1473—

Instruction Following Not comparable

ERNIE 5.1: 76.7 (#37), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkERNIE 5.1Qwen3 8B
LMArena Instruction Following1460—

Long Context ERNIE 5.1 leads

ERNIE 5.1: 44.7 (#59), Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkERNIE 5.1Qwen3 8B
Fiction.LiveBench—62.1%
LMArena Longer Query1462—

Writing & Preference Not comparable

ERNIE 5.1: 65.1 (#52), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkERNIE 5.1Qwen3 8B
LMArena Text1468—
LMArena Creative Writing1441—
LMArena Multi-Turn1471—

Frequently asked questions

Is ERNIE 5.1 better than Qwen3 8B?

ERNIE 5.1 is the stronger model overall, scoring 43.8 to 33.7 on the Noometry Index.

Is ERNIE 5.1 or Qwen3 8B better for coding?

ERNIE 5.1 scores higher on coding benchmarks: 44.0 versus 34.0 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Qwen3 8B share?

0 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper