Model comparison

Claude 3.5 Sonnet vs Llama 13b

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 24.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 5 categories and Llama 13b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude 3.5 Sonnet leads 52.9 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Llama 13b specifications
Claude 3.5 SonnetLlama 13b
ProviderAnthropicMeta
Noometry Index34.624.4
Released2024-06-202023-02-24
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6021

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Coding1342683
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Llama 13b: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Hard Prompts1305728
Epoch Capabilities Index133.55100.58
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
BIG-Bench Hard—37.9%
ForecastBench60.7—
HellaSwag—79.2%
LAMBADA—75.2%
LiveBench59%—
PIQA—80.1%
WinoGrande—73%

Math Llama 13b leads

Claude 3.5 Sonnet: 19.2 (#288), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Math1307838
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—
GSM8K—20.6%

Knowledge Not comparable

Claude 3.5 Sonnet: 28.6 (#245), Llama 13b: —

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
MMLU87.3%47.7%
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
LMArena Expert1265—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Llama 13b: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—
ScienceQA—43.3%

Multilingual Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 43.2 (#185), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Non-English1283819
LMArena Chinese1272—
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Russian1306—
LMArena Spanish1290—

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Instruction Following1297781
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Not comparable

Claude 3.5 Sonnet: 39.9 (#167), Llama 13b: —

Long Context benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Longer Query1311—

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetLlama 13b
LMArena Text1298834
LMArena Creative Writing1292794
LMArena Multi-Turn1326753
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Llama 13b?

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 24.4 on the Noometry Index.

Is Claude 3.5 Sonnet or Llama 13b better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 21.4 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Llama 13b share?

10 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper