Model comparison

Llama 3.2 1B vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 20.1 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 3.2 1B scores higher in 0 categories and Mercury in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury leads 46.2 to 21.3.
  • Llama 3.2 1B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 1B and Mercury specifications
Llama 3.2 1BMercury
ProviderMetaInception
Noometry Index20.137.6
Released2024-09-24—
WeightsOpenProprietary
Context window60K—
Max output54K—
Input $ / M tokens$0.027—
Output $ / M tokens$0.20—
Results tracked229

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Llama 3.2 1B: 21.1 (#338), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkLlama 3.2 1BMercury
LMArena Coding10701322
BigCodeBench Instruct8.2%—
BigCodeBench Complete11.3%—

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BMercury
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Mercury leads

Llama 3.2 1B: 16.2 (#308), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkLlama 3.2 1BMercury
LMArena Hard Prompts10441285
Kagi LLM Benchmark—21.6%
Chess Puzzles0%—
Epoch Capabilities Index101.99—

Math Not comparable

Llama 3.2 1B: 10.4 (#313), Mercury: —

Math benchmarks
BenchmarkLlama 3.2 1BMercury
OTIS Mock AIME 2024-20250.6%—
LMArena Math1086—

Knowledge Not comparable

Llama 3.2 1B: 7.2 (#312), Mercury: —

Knowledge benchmarks
BenchmarkLlama 3.2 1BMercury
GPQA Diamond23.9%—
LMArena Expert1007—

Multilingual Mercury leads

Llama 3.2 1B: 23.8 (#292), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkLlama 3.2 1BMercury
LMArena Non-English9731260
LMArena Chinese959—
LMArena German1014—
LMArena Russian941—

Instruction Following Mercury leads

Llama 3.2 1B: 52.4 (#290), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BMercury
LMArena Instruction Following10311239

Long Context Mercury leads

Llama 3.2 1B: 31.9 (#274), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkLlama 3.2 1BMercury
LMArena Longer Query10501266

Writing & Preference Mercury leads

Llama 3.2 1B: 21.3 (#310), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BMercury
LMArena Text10551282
LMArena Creative Writing10331191
LMArena Multi-Turn10301282
EQ-Bench Creative Writing200—

Frequently asked questions

Is Llama 3.2 1B better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 20.1 on the Noometry Index.

Is Llama 3.2 1B or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 21.1 in the Noometry coding category.

How many benchmarks do Llama 3.2 1B and Mercury share?

8 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper