Model comparison

Llama 3-8B vs Sonar

Sonar is the stronger model overall, scoring 38.5 to 25.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Sonar Perplexity

38.5

Rank #187 Confirmed

Summary

  • The widest gap is in math, where Sonar leads 33.7 to 8.8.
  • Llama 3-8B has downloadable open weights; the other is API-only.

Side by side

Llama 3-8B and Sonar specifications
Llama 3-8BSonar
ProviderMetaPerplexity
Noometry Index25.538.5
Released2024-04-182024-01-01
WeightsOpenProprietary
Context window—128K
Max output—4K
Input $ / M tokens—$1
Output $ / M tokens—$1
Results tracked347

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Sonar leads

Llama 3-8B: 31.0 (#289), Sonar: 35.7 (#221)

Coding benchmarks
BenchmarkLlama 3-8BSonar
BigCodeBench Instruct31.9%—
LiveBench Coding—35.1%
LMArena Coding1152—
BigCodeBench Complete36.9%—
HumanEval+56.7%—
MBPP+54.8%—

Reasoning Sonar leads

Llama 3-8B: 14.3 (#326), Sonar: 21.1 (#227)

Reasoning benchmarks
BenchmarkLlama 3-8BSonar
Chess Puzzles0%—
LiveBench Reasoning—46.3%
LMArena Hard Prompts1133—
DTBench43.9%—
LiveBench Data Analysis—37.9%
Adversarial NLI57.3%—
Epoch Capabilities Index116.45—
ForecastBench58.6—
LiveBench—46.9%
WinoGrande75.7%—

Math Sonar leads

Llama 3-8B: 8.8 (#323), Sonar: 33.7 (#200)

Math benchmarks
BenchmarkLlama 3-8BSonar
OTIS Mock AIME 2024-20251.9%—
LiveBench Math—41.6%
LMArena Math1151—
MATH Level 56.1%—

Knowledge Not comparable

Llama 3-8B: 7.8 (#308), Sonar: —

Knowledge benchmarks
BenchmarkLlama 3-8BSonar
GPQA Diamond26.1%—
LMArena Expert1113—
ARC (AI2) Challenge82.8%—
MMLU68.8%—
OpenBookQA82.6%—
TriviaQA67.7%—

Multilingual Not comparable

Llama 3-8B: 30.8 (#261), Sonar: —

Multilingual benchmarks
BenchmarkLlama 3-8BSonar
LMArena Non-English1098—
LMArena Chinese1076—
LMArena French1159—
LMArena German1104—
LMArena Japanese967—
LMArena Korean1004—
LMArena Russian1109—
LMArena Spanish1173—

Instruction Following Sonar leads

Llama 3-8B: 58.4 (#260), Sonar: 71.4 (#150)

Instruction Following benchmarks
BenchmarkLlama 3-8BSonar
LiveBench Instruction Following—76.2%
LMArena Instruction Following1127—

Long Context Not comparable

Llama 3-8B: 34.2 (#251), Sonar: —

Long Context benchmarks
BenchmarkLlama 3-8BSonar
LMArena Longer Query1128—

Writing & Preference Sonar leads

Llama 3-8B: 37.5 (#256), Sonar: 52.6 (#167)

Writing & Preference benchmarks
BenchmarkLlama 3-8BSonar
LMArena Text1166—
LMArena Creative Writing1150—
LMArena Multi-Turn1152—
LiveBench Language—44.1%

Frequently asked questions

Is Llama 3-8B better than Sonar?

Sonar is the stronger model overall, scoring 38.5 to 25.5 on the Noometry Index.

Is Llama 3-8B or Sonar better for coding?

Sonar scores higher on coding benchmarks: 35.7 versus 31.0 in the Noometry coding category.

How many benchmarks do Llama 3-8B and Sonar share?

0 benchmarks have published results for both models. Llama 3-8B has 34 scored results on Noometry and Sonar has 7.

Related comparisons

Go deeper