Model comparison

Llama 3-70B vs Yi-34B

Llama 3-70B and Yi-34B score almost the same on the Noometry Index (28.8 vs 27.8), so choose on price, context window or the category you care about most.

Last verified . 21 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Yi-34B 01.AI

27.8

Rank #329 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Llama 3-70B scores higher in 6 categories and Yi-34B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3-70B leads 20.8 to 7.5.
  • The biggest single-benchmark swing is GPQA Diamond: 40.6% for Llama 3-70B and 14.7% for Yi-34B.

Side by side

Llama 3-70B and Yi-34B specifications
Llama 3-70BYi-34B
ProviderMeta01.AI
Noometry Index28.827.8
Released2024-04-182023-11-02
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3123

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

Llama 3-70B: 35.8 (#218), Yi-34B: 32.3 (#274)

Coding benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Coding12061112
BigCodeBench Instruct43.6%—
BigCodeBench Complete54.5%—
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use Not comparable

Llama 3-70B: 21.1 (#139), Yi-34B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BYi-34B
Cybench5%—

Reasoning Yi-34B leads

Llama 3-70B: 18.0 (#288), Yi-34B: 21.2 (#226)

Reasoning benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Hard Prompts11951104
Epoch Capabilities Index122.93117.39
Kagi LLM Benchmark35.1%—
DTBench54.2%—
BIG-Bench Hard—71.7%
ForecastBench57.1—
WinoGrande83.5%—

Math Yi-34B leads

Llama 3-70B: 12.8 (#305), Yi-34B: 21.6 (#282)

Math benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Math12181114
MATH Level 522.6%5.1%
OTIS Mock AIME 2024-20254.3%—
GSM8K—76%

Knowledge Llama 3-70B leads

Llama 3-70B: 20.8 (#277), Yi-34B: 7.5 (#309)

Knowledge benchmarks
BenchmarkLlama 3-70BYi-34B
GPQA Diamond40.6%14.7%
LMArena Expert11491061
MMLU79.3%76.3%

Multilingual Llama 3-70B leads

Llama 3-70B: 33.6 (#251), Yi-34B: 29.7 (#264)

Multilingual benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Non-English11421079
LMArena Chinese11141176
LMArena French12321081
LMArena German11691042
LMArena Japanese1017993
LMArena Korean1017959
LMArena Russian11591050
LMArena Spanish12411070

Instruction Following Llama 3-70B leads

Llama 3-70B: 62.5 (#238), Yi-34B: 56.2 (#274)

Instruction Following benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Instruction Following11941091

Long Context Llama 3-70B leads

Llama 3-70B: 35.6 (#240), Yi-34B: 33.2 (#264)

Long Context benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Longer Query11741094

Writing & Preference Llama 3-70B leads

Llama 3-70B: 42.8 (#231), Yi-34B: 34.1 (#273)

Writing & Preference benchmarks
BenchmarkLlama 3-70BYi-34B
LMArena Text12211129
LMArena Creative Writing12101108
LMArena Multi-Turn12231113

Frequently asked questions

Is Llama 3-70B better than Yi-34B?

Llama 3-70B and Yi-34B score almost the same on the Noometry Index (28.8 vs 27.8), so choose on price, context window or the category you care about most.

Is Llama 3-70B or Yi-34B better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 32.3 in the Noometry coding category.

How many benchmarks do Llama 3-70B and Yi-34B share?

21 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and Yi-34B has 23.

Related comparisons

Go deeper