Model comparison

Llama 13b vs Mercury 2.5

Mercury 2.5 is the stronger model overall, scoring 33.5 to 24.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • The widest gap is in coding, where Mercury 2.5 leads 39.5 to 21.4.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Mercury 2.5 specifications
Llama 13bMercury 2.5
ProviderMetaInception
Noometry Index24.433.5
Released2023-02-242026-09-08
WeightsOpenProprietary
Context window—260K
Max output—66K
Input $ / M tokens—$0.04
Output $ / M tokens—$0.15
Results tracked214

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Llama 13b: 21.4 (#337), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkLlama 13bMercury 2.5
SciCode—38.5%
LMArena Coding683—
ALE-Bench—301.65

Reasoning Mercury 2.5 leads

Llama 13b: 14.0 (#329), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkLlama 13bMercury 2.5
CritPt—0%
LMArena Hard Prompts728—
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkLlama 13bMercury 2.5
ProofBench—3%
LMArena Math838—
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Mercury 2.5: —

Knowledge benchmarks
BenchmarkLlama 13bMercury 2.5
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Mercury 2.5: —

Multimodal benchmarks
BenchmarkLlama 13bMercury 2.5
ScienceQA43.3%—

Multilingual Not comparable

Llama 13b: 16.6 (#297), Mercury 2.5: —

Multilingual benchmarks
BenchmarkLlama 13bMercury 2.5
LMArena Non-English819—

Instruction Following Not comparable

Llama 13b: 36.7 (#305), Mercury 2.5: —

Instruction Following benchmarks
BenchmarkLlama 13bMercury 2.5
LMArena Instruction Following781—

Writing & Preference Not comparable

Llama 13b: 13.8 (#312), Mercury 2.5: —

Writing & Preference benchmarks
BenchmarkLlama 13bMercury 2.5
LMArena Text834—
LMArena Creative Writing794—
LMArena Multi-Turn753—

Frequently asked questions

Is Llama 13b better than Mercury 2.5?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 24.4 on the Noometry Index.

Is Llama 13b or Mercury 2.5 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Mercury 2.5 share?

0 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper